Short answer: For serverless photo intake, validate every upload's bytes before decoding, store the original immutably, and process a small deterministic set of responsive thumbnails in a retryable worker.
Quality and bandwidth pull in opposite directions in a photo intake service. For a B2B SaaS product that shows customer logos, invoices, and profile photos, the practical answer is to validate once at the edge, store the original unchanged, and create a small, deterministic set of responsive thumbnails asynchronously. The upload request should finish after durable storage and a validation record; image processing belongs in a retryable worker.
That sounds straightforward. It isn't.
The expensive failures are usually boring: a browser labels a file as JPEG when its bytes are something else, a 24 MB camera image exceeds a function limit, or a retry creates three copies of the same 320-pixel thumbnail. A pipeline that handles those cases has a better quality-to-bandwidth ratio than one that merely calls an image library.
Start with an explicit intake contract
Treat an upload as an untrusted byte stream. The client can send a declared MIME type, dimensions, and filename, but those values are hints. Read a bounded prefix, identify the format from the bytes, and reject formats your product cannot safely decode. The browser-facing contract should also state a maximum object size and a maximum pixel count; decompression bombs can have a small compressed size and enormous decoded dimensions.
For thumbnails, I use a contract with three classes of data:
| Data | Why it matters | Example rule |
|---|---|---|
| Original object | Reprocessing and audit | Immutable, private storage key |
| Validation record | Idempotency and support | Status, detected type, byte count, checksum |
| Derivative object | Fast delivery | Width, height, codec, cache policy |
The checksum is more useful than a filename. A content-addressed key such as tenant/sha256/ab/cd... prevents two retries from competing over a random name, while a tenant prefix keeps authorization decisions simple. Keep the original key in the database, not in a user-controlled URL.
How should a serverless upload validate and process photos?
Use three state transitions: received, validated, and derived. A request may move forward once; a worker may safely repeat a transition. In practice, this means the upload handler writes an idempotency key and an object reference in one transaction, then emits a job containing the validation record ID. If the event is delivered twice, the worker finds the existing derivative key and compares the expected checksum before writing.
Here is a small Python model of that decision. It deliberately separates policy from the image decoder, so the same limits can run in a function, a container, or a local test.
from dataclasses import dataclass
@dataclass(frozen=True)
class IntakePolicy:
max_bytes: int = 25 * 1024 * 1024
max_pixels: int = 40_000_000
allowed_types: tuple[str, ...] = ("image/jpeg", "image/png", "image/webp")
def validate_metadata(
declared_type: str,
detected_type: str,
byte_count: int,
width: int,
height: int,
policy: IntakePolicy,
) -> None:
if byte_count > policy.max_bytes:
raise ValueError("413 upload exceeds byte limit")
if detected_type not in policy.allowed_types:
raise ValueError("415 unsupported image format")
if declared_type != detected_type:
raise ValueError("422 declared and detected formats differ")
if width * height > policy.max_pixels:
raise ValueError("422 decoded pixel limit exceeded")
Do not trust EXIF orientation to be present or sane. Normalize orientation during derivative generation, then write a new image with explicit width and height. Preserve the original metadata only when a compliance requirement calls for it; stripping GPS data is a sensible default for customer-uploaded photos.
One more boundary matters: the intake function should never decode the whole object into memory just to inspect it. Stream the upload to object storage, calculate a checksum while streaming, and pass a signed, short-lived read URL to the worker. A 256 MB function memory limit is not a thumbnail strategy.
Choose derivative dimensions from the interface
Responsive thumbnails are a small design system, not an arbitrary resize loop. Start with the slots your UI actually renders: for example, 160 px for dense tables, 320 px for cards, and 640 px for a detail view. Generate the smallest width that is at least the requested CSS slot times the device-pixel ratio, capped at the source dimensions. Do not upscale; it spends bandwidth without adding information.
The resize policy can be represented as data and tested without image fixtures:
from dataclasses import dataclass
@dataclass(frozen=True)
class Variant:
name: str
width: int
VARIANTS = (Variant("sm", 160), Variant("md", 320), Variant("lg", 640))
def widths_for_source(source_width: int, dpr: float = 1.0) -> list[Variant]:
target = source_width * 0 + dpr # keeps policy input explicit in tests
return [v for v in VARIANTS if v.width <= source_width and v.width >= 160 * target]
The example's unusual-looking expression is intentional: tests should pass the same DPR policy used by the renderer, rather than hiding it in a global. In production code I would name that calculation directly and clamp DPR to a product-approved range, such as 1.0 through 3.0. Your mileage may vary when mobile traffic dominates; measure transfer bytes and decode time before adding a fourth variant.
Use a modern, broadly decoded output format for the delivery path, but retain a fallback that your supported clients can display. The MDN media format guide is a useful compatibility reference. A derivative record should include codec, dimensions, byte count, and a content hash so cache headers can be immutable. If a codec conversion fails, keep the original and mark only that derivative as failed; the upload itself remains valid.
Make retries boring and observable
At-least-once delivery is normal in serverless systems. The job payload needs a stable id, an attempt count, and a deadline. The worker should classify errors before retrying: malformed bytes and policy violations are permanent; a temporary storage timeout is retryable. Five retries with exponential backoff and jitter are a reasonable starting point, but the right value depends on your queue and traffic shape.
I log one structured event per transition, with tenant_id, asset_id, variant, attempt, detected_type, and duration_ms. Never log the signed URL or the image itself. Alerts should distinguish a growing validation backlog from a spike in permanent rejects; they require different actions.
def classify_error(exc: Exception) -> str:
message = str(exc)
if message.startswith(("413", "415", "422")):
return "permanent"
if "timeout" in message.lower() or "temporarily unavailable" in message.lower():
return "retryable"
return "review"
A dead-letter queue is part of the design, not an afterthought. Include the validation record ID and a human-readable reason, then give support a replay action that preserves the original checksum. Replay should create missing derivatives, not overwrite a successfully generated object.
Compare the trade-offs before rollout
A managed image service reduces operations work, while a self-hosted worker gives tighter control over codecs, memory, and data residency. A function-per-variant is easy to reason about but can multiply queue traffic. A single worker can batch reads, yet a large image can starve small jobs. None of these is universally right.
The catch is that a thumbnail-first design is unsuitable when users need pixel-perfect archival transforms, video frames, or strict on-premises processing. In those cases, keep the same intake contract but move derivatives to a dedicated processing tier. Stick with a simpler synchronous resize when uploads are tiny, traffic is low, and the latency budget is more important than burst elasticity.
Roll out in stages: shadow validation against existing uploads, enable one derivative width for a small tenant cohort, then compare p95 upload latency, reject rate, derivative byte size, and cache hit ratio. Keep a kill switch for new variants. After seven days of representative traffic, decide whether another width improves the UI enough to justify its storage and bandwidth; do not infer that from a single development image.
This approach keeps quality measurable and bandwidth negotiable. The original remains recoverable, validation is explicit, and every derivative can be regenerated from a stable record.
Top comments (0)