DEV Community

DemetriusReed2163
DemetriusReed2163

Posted on

E-commerce Thumbnail Intake: Choosing Upload-Time Processing with Durable Asset Identity

Short answer: process thumbnails at upload when the storefront needs predictable first-view latency, but make the upload an idempotent job with a durable asset ID; choose on-demand derivatives when most images are rarely viewed or when the rendition matrix changes often.

That decision is easy to phrase and hard to measure. A React uploader can show a progress bar in seconds, yet the catalog page may still be waiting on a missing 320-pixel image, a browser that sent an untrusted MIME type, or a worker retry that created a second object. I build the measurement around those failure modes before choosing a default.

The experiment is deliberately small: take a week of representative product photos, replay uploads with the same concurrency as a peak seller import, and compare time-to-first-thumbnail, derivative hit rate, bytes written, and reprocessing volume. A notebook is useful for exploring the distribution; production needs the same counters emitted by the worker. The failed/simple approach is to resize in the browser and upload whatever it produces. It saves a request, but it makes dimensions, codecs, and trust decisions client-controlled. That is a poor contract for a marketplace. In one replay, I would deliberately include a 12 MB phone photo, a PNG whose extension says JPEG, a seller who double-clicks the submit button, and a worker restart after the source write but before the ready transition. The point is not to manufacture a scary demo; it is to force every boundary to answer the same question: which asset ID owns these bytes, which recipe produced this derivative, and can the next attempt prove it is repeating work rather than inventing a new object? The resulting trace should let an engineer follow one upload from browser event to catalog read without opening the object store and guessing.

Measure twice.

What should browser image uploads guarantee before a thumbnail exists?

The server should assign an opaque asset ID as soon as it accepts a complete upload. The ID is the stable reference used by the React app, catalog rows, moderation records, and derivative jobs. Object keys can change; the asset ID should not. Persist the source byte count, measured dimensions, content type, checksum, tenant, and processing state beside it. Do not infer ownership from a filename supplied by the browser.

I use a small state machine: received, processing, ready, and rejected. A duplicate request with the same idempotency key returns the original asset ID instead of creating another source object. A worker can safely retry processing because each derivative key includes the asset ID, width, and recipe version. This is where a five-line shortcut becomes a support ticket.

from dataclasses import dataclass
from hashlib import sha256


@dataclass(frozen=True)
class Upload:
    asset_id: str
    tenant_id: str
    content_type: str
    byte_count: int
    width: int
    height: int
    checksum: str


def validate_upload(asset_id: str, tenant_id: str, body: bytes,
                    declared_type: str, width: int, height: int) -> Upload:
    allowed = {"image/jpeg", "image/png", "image/webp"}
    if declared_type not in allowed:
        raise ValueError("unsupported image type")
    if not body or width <= 0 or height <= 0:
        raise ValueError("empty or unreadable image")
    digest = sha256(body).hexdigest()
    return Upload(asset_id, tenant_id, declared_type, len(body), width, height, digest)


def derivative_key(upload: Upload, width: int, recipe: str) -> str:
    return f"{upload.tenant_id}/{upload.asset_id}/{recipe}/{width}.webp"
Enter fullscreen mode Exit fullscreen mode

The declared type is only a hint; a decoder must confirm the actual bytes and dimensions. MDN's media format guidance is a useful compatibility reference, but your supported browser matrix is the final test. Keep the original immutable, and record the measured metadata rather than trusting a React File property.

Keep the source boring.

How can upload-time processing balance latency, storage, and responsive thumbnails?

Upload-time work buys a predictable read path. When a seller opens a listing immediately after upload, the 320, 640, and 1280 variants are already scheduled, and the page does not pay a cold transformation penalty. The cost is write amplification: every image may produce three objects even if nobody opens the detail page.

On-demand work reverses that trade. It keeps storage lean for long-tail catalog items, but the first viewer pays queue and transform latency. It also needs a cache stampede policy. Ten simultaneous product-page requests should converge on one derivative job, not enqueue ten jobs with identical inputs.

I would start with upload-time generation for the hot path and reserve on-demand generation for uncommon widths or older catalog rows. Measure p95 first-thumbnail latency and the percentage of generated variants never read after 30 days. If the unused ratio dominates, move that width to on-demand. Your mileage may vary by catalog shape; a seasonal seller import behaves differently from a steady stream of single-item uploads.

A Python worker that stays idempotent

The worker consumes an asset ID, loads the immutable source, and writes each derivative under a deterministic key. A database uniqueness constraint on (asset_id, recipe, width) is the final duplicate guard. Queue retries should preserve the same job ID, and a timeout should leave the state recoverable rather than guessing whether bytes were written.

from typing import Protocol


class Store(Protocol):
    def read(self, key: str) -> bytes: ...
    def write_if_absent(self, key: str, data: bytes, content_type: str) -> bool: ...


def build_derivatives(store: Store, source_key: str, upload: Upload,
                      widths: tuple[int, ...] = (320, 640, 1280)) -> list[str]:
    source = store.read(source_key)
    written = []
    for width in widths:
        key = derivative_key(upload, width, "thumb-v1")
        data = resize_and_encode(source, width, "image/webp")
        store.write_if_absent(key, data, "image/webp")
        written.append(key)
    return written
Enter fullscreen mode Exit fullscreen mode

The resize function is intentionally an adapter boundary: the image library can change without changing asset identity or queue semantics. Keep decode limits, maximum pixel area, and output quality in configuration, then include the recipe version in every key. A later recipe creates a new derivative; it does not overwrite one that a cached page may still reference.

Which evidence should decide the processing strategy?

An eval harness should replay real dimensions and formats, then score four things separately: visual acceptability, first-view latency, bytes written, and duplicate-job rate. Do not hide image quality inside a single average. A thumbnail that is fast but crops a product label incorrectly is a failed catalog experience.

I once assumed a successful HTTP response meant the pipeline was done. It wasn't. The response only acknowledged intake; the durable contract was the asset row plus observable state transitions. Emit counters for queue age, decode rejection, derivative completion, cache hits, and retries. Store opaque IDs in logs, not signed URLs or customer filenames.

The catch is that upload-time processing is not suitable when storage must remain minimal, when users can request an unbounded rendition matrix, or when source images are routinely replaced before anyone views them. Stick with on-demand derivatives in those cases, and add a single-flight lock so concurrent readers share one job. Conversely, on-demand is a poor fit for a storefront promise that every newly published listing must render instantly on a slow mobile connection.

Start with one recipe and three widths, run the replay, and make the threshold explicit. I am not sure which side will win for your catalog until those counters exist. That uncertainty belongs in the design, not in a hidden default.

References

Top comments (1)