DEV Community

DemetriusReed2163
DemetriusReed2163

Posted on

Catalog Transformation Presets: Python Manifests Beat Mutable URLs for OCR Derivatives

Short answer: define each catalog image transformation preset as an immutable, versioned manifest, then derive the output object key from the source digest, preset version, and encoder policy. Mutable processing URLs are fine for exploration, but a manifest wins once OCR results, cache behavior, and storage cost have to be repeatable.

The evaluation constraint matters more than the resize syntax. For a headless-commerce catalog, a derivative is successful only if it fits the display slot, preserves the text a shopper or downstream model needs, and reuses an existing object when the inputs have not changed. A notebook can prove that one label photo looks sharp. Production has to prove that the same recipe handles portrait bottles, wide cartons, rotated care tags, and a revised source image without quietly multiplying stored files.

This is the choice: use URL parameters while discovering useful dimensions and crops; promote a recipe into a versioned Python manifest before other services depend on it. The catch is that manifests add review and migration work. For a tiny, disposable catalog with no OCR regression suite and no shared cache, stick with parameterized URLs until the recipe stabilizes.

How should catalog teams create and reuse image transformation presets?

Treat a preset as data, not as a clever string builder. It needs a stable name for humans, a version for migrations, geometry rules, an output media type, an encoder policy, and an explicit statement of intent. The intent is important: product_card and ocr_detail may happen to share dimensions today, yet they optimize for different outcomes and should be free to diverge later.

A useful boundary is to keep selection separate from execution. Catalog code asks for ocr_detail@3; an image worker resolves that identifier, reads the source object, performs the transform, and writes the result at a deterministic key. The storefront never reconstructs encoder options. The OCR pipeline doesn't invent a second crop rule. One contract travels from notebook to production — small enough to inspect, strict enough to evaluate.

Don't overwrite version 3 in place.

An edit that changes pixels is a new version, even if the preset's friendly name stays put. That rule prevents a warm cache from serving yesterday's image under today's semantics, and it gives an eval run an exact configuration to report. A mutable URL such as ?width=1200&quality=80 looks transparent, but transparency isn't identity: defaults can live outside the URL, parameter order can vary, and two callers can encode the same idea differently. A canonical manifest turns those loose choices into one reviewable record.

The format decision belongs in that record as well. Browser and container support differ by image format, so the output policy should be chosen against the clients that consume it rather than copied from a generic optimization checklist. The MDN media format guide is a practical compatibility reference for that decision.

A focused Python manifest and cache key

The following example catalogs two intents without implementing an image codec. That separation is deliberate. A real worker can map the validated fields onto its chosen processor, while the identity and reuse rules remain ordinary Python.

from dataclasses import asdict, dataclass
from hashlib import sha256
import json


@dataclass(frozen=True)
class Preset:
    name: str
    version: int
    width: int
    height: int
    fit: str
    media_type: str
    encoder_quality: int
    purpose: str

    def canonical_bytes(self) -> bytes:
        payload = json.dumps(
            asdict(self),
            sort_keys=True,
            separators=(",", ":"),
        )
        return payload.encode("utf-8")


PRESETS = {
    "product_card": Preset(
        name="product_card",
        version=2,
        width=640,
        height=800,
        fit="cover",
        media_type="image/webp",
        encoder_quality=82,
        purpose="consistent catalog grid",
    ),
    "ocr_detail": Preset(
        name="ocr_detail",
        version=3,
        width=1600,
        height=1600,
        fit="contain",
        media_type="image/png",
        encoder_quality=100,
        purpose="retain packaging text for OCR",
    ),
}


def derivative_key(
    source_bytes: bytes,
    preset: Preset,
    extension: str,
) -> str:
    source_digest = sha256(source_bytes).hexdigest()
    preset_digest = sha256(preset.canonical_bytes()).hexdigest()[:16]
    return (
        f"derivatives/{source_digest}/"
        f"{preset.name}/v{preset.version}-{preset_digest}.{extension}"
    )


source = b"example catalog photo bytes"
key = derivative_key(source, PRESETS["ocr_detail"], "png")
assert key == derivative_key(source, PRESETS["ocr_detail"], "png")
Enter fullscreen mode Exit fullscreen mode

Including both the version and a digest may look redundant. It pays for itself during review: the version is readable in logs, while the digest catches an accidental field change that did not receive a version bump. A deployment check can serialize every manifest and reject a duplicate (name, version) with a different digest. That's a policy failure, not an image-processing mystery.

The full source digest also answers a common catalog problem. If a merchant replaces a photo while retaining the SKU and filename, the derivative key still changes. If the source bytes and manifest stay identical, workers converge on the same key and can reuse the stored object. No timestamp is needed, and a retry doesn't mint a fresh derivative.

There is one operational nuance: creation needs idempotent write semantics. Two workers may miss the same cache at nearly the same moment. Both can compute the same bytes and target the same deterministic key; the storage layer should allow a conditional create or accept an identical replacement according to the team's consistency policy. The manifest does not remove concurrency. It gives concurrency one destination.

Evaluate text retention before promoting a recipe

For OCR, pixel dimensions are an input, not the score. Build a small, frozen evaluation set that represents the catalog's awkward material: fine-print ingredient panels, glossy labels, vertical text, low-contrast embossing, and photos where the product occupies only part of the frame. Keep the source asset identifiers and expected text assertions beside the preset version. Then run the candidate and current preset over the same set.

I wouldn't promote a candidate from visual inspection alone. The decision table should combine OCR quality with storage and reuse signals, because a beautiful output that doubles the number of retained objects can still be the wrong catalog default. I'm not sure what OCR threshold fits every catalog; language mix, packaging, and the cost of a missed token change the acceptable floor. A team resolves that uncertainty with labeled examples and an explicit acceptance rule, not a universal number.

Signal What to record Promotion question
Text recall Expected strings found per fixture Does the candidate avoid regressions on required label text?
Crop survival Required regions retained Are barcodes, warnings, and edge text still present?
Encoded bytes Bytes per generated object Is the storage change justified by measured OCR gain?
Reuse ratio Requests served from an existing derivative key Are callers actually sharing the preset?
Variant count Stored derivatives per source digest Did a caller create an unplanned combination?
Processing time Duration by preset version Can the worker meet its queue budget?

One failed fixture should be legible in the report. Consider fixture=carton_side_017: the expected assertion includes a warning printed near the right fold, while the candidate uses a tighter center crop. Record preset=ocr_detail@4, the expected phrase, the extracted phrase, crop bounds, encoded bytes, and the prior preset's result in one row. If the phrase disappears, the team can distinguish a crop-policy regression from an OCR-model change and reject version 4 before it becomes a shared catalog contract. If the phrase survives but encoded size rises, reviewers can decide whether that particular gain earns its storage cost. This is far more actionable than a single average score, which can hide a required warning behind many easy front-label examples. It also connects prompt and token cost to image work: when OCR text feeds a model, log the extracted character count and downstream token count by preset version. A crop change can alter inference cost even when the image file gets smaller, so the eval artifact needs to keep image, OCR, and model-input measurements tied to the same fixture and manifest digest.

The mismatch matters.

Keep the notebook honest. It should import the same manifest schema as the worker, emit the candidate's canonical JSON, and produce an eval artifact that CI can compare. The notebook is for exploration; the checked-in manifest and fixtures are the production claim.

Storage and cache economics decide the rollout

Count objects before chasing byte-level savings. An uncontrolled parameter space can create many near-duplicate derivatives for one source: widths of 639 and 640, equivalent crop spellings, or several quality values selected by different clients. A named preset collapses those choices. Track unknown preset requests as contract violations, and allow only registered versions at the worker boundary.

Then measure bytes.

Keep both counts.

The useful cost model covers retained source bytes, retained derivative bytes, write operations, reads or delivery, invalidation, and compute for cache misses. Exact billing units depend on the selected infrastructure, so a portable design should emit quantities rather than hard-code currency into application logic. Record source_digest, preset_name, preset_version, derivative_bytes, cache_outcome, and duration_ms for each generation attempt. Aggregate by preset and catalog cohort. This makes a migration observable without tying the schema to one provider's invoice.

Roll out a new version with dual reads and single writes: request the new key, generate it if absent, and keep the prior version addressable during a defined comparison window. Don't eagerly regenerate an entire catalog unless the access pattern or business deadline justifies it. Lazy creation limits speculative storage, while a scheduled backfill is better when a launch requires predictable first-request latency. Neither policy wins everywhere — the catalog's hot set and update cadence decide.

The manifest approach is not suitable when every transformation is genuinely bespoke, such as an interactive editor where a user controls arbitrary crop coordinates and expects a one-off export. In that case, preserve a canonical operation document per edit rather than forcing thousands of user choices into named presets. Mutable URLs also remain reasonable for local experiments that will never become shared contracts. Once multiple services reuse an output, though, versioned identity is the cleaner boundary.

Before copying this design, measure four things on your own fixtures: OCR regression by required text, derivative bytes per source, reuse ratio by preset version, and the number of unregistered parameter combinations. Promote the manifest only when it improves repeatability without violating the storage budget. That is the production threshold, not the fact that one transformed photo looked good in a notebook.

References

Further reading

Top comments (0)