Keep the original image file. Not the finished cutout, not the 2048 px file you were about to call a master — the raw upload, plus the alpha mask that background removal produced. Pick a ladder of derivative sizes for the layout you ship today, render it, and treat every rendered file as a cache entry you're allowed to lose. The constraint that decides the whole design is a byte budget per product card, not how pretty one cutout looks at full resolution.
I assumed the cutout was the master. A design refresh corrected me.
The system here is the catalog for a developer-tools store: dev boards, probes, cables, a few keyboards, photographed against whatever the vendor sent us, then cut out so every listing card shows the product on one uniform canvas. Quality against bandwidth is the only axis that matters on those cards, because a grid of thirty products is thirty image requests on a phone. When the layout changes, the ladder changes with it, and every byte you already shipped is suddenly the wrong size.
The cutout master that couldn't survive a layout change
The first version of this pipeline was as simple as it gets. Segment the product, composite it on white, save a 2048 px lossy file with alpha, throw the camera original away, and derive 320 / 640 / 1024 from that master on demand. Storage stayed small. Every card looked fine.
Then the refresh landed: wider cards, a dark canvas behind them, and a new 1536 px slot on the product page that the old master couldn't fill without upscaling. Re-encoding a lossy file into another lossy file compounds what the first encoder already threw away, and the damage concentrates exactly where a cutout lives — the high-contrast boundary between product and background. Worse, we had flattened onto white, so on the dark canvas every product wore a pale fringe along its edge, which is what you get when partially transparent edge pixels were blended against one background color and then displayed over another. None of that was recoverable from the derivative. The pixels were gone.
That's the part worth internalizing: a derivative encodes two decisions — the size and the surrounding canvas — and a design refresh is precisely the event that invalidates both.
What do you keep when a design refresh changes every derivative size?
Three artifacts, in descending order of cost to recreate.
The original upload, stored immutably and addressed by the digest of its bytes, in a cold tier. It's the only thing in the system nobody can reconstruct — reshooting a discontinued SKU isn't a budget line, it's a phone call to a vendor who no longer stocks it.
The mask, as a single-channel lossless file, next to the recipe that produced it: model id, model revision, feather radius, the crop box a human approved, and whatever trim threshold you settled on. This is the artifact most pipelines skip, and skipping it is expensive in a way that shows up on an inference bill rather than a storage bill. Background removal is the one step that costs real money per call and isn't deterministic across model versions; compositing is arithmetic. Keep the mask and a full catalog reprocess is a compositing job you can run on a laptop over lunch. Throw it away and the same refresh becomes a re-segmentation run over every SKU, at inference prices, with a fresh chance of regressions on the products that were hard to cut in the first place. Cables and mesh grilles, mostly. Anything with hair-like edges.
Then the manifest: which derivative keys exist for which SKU, which encoder revision wrote them, and which ladder they belong to. That's what lets you flip a whole catalog to a new generation atomically instead of purging your way through it.
I'm not sure mask reuse is always the right call across a model upgrade. When the segmentation model itself improves, you probably want new masks, not old ones replayed — so pin the model revision in the recipe and treat a model bump as a genuine reprocess with its own review, rather than a silent recomposite.
Quality against bandwidth, decided per slot
Per-slot budgets are what keep this honest. A card thumbnail and a product hero have nothing in common except the source pixels, so giving them one global quality number guarantees you overspend on one and underspend on the other.
| Slot | Rendered width | Alpha needed | Budget per file | What failure looks like |
|---|---|---|---|---|
| Grid card | 384 px | no, canvas baked in | 24 KB | mushy logo silkscreen on a PCB |
| Card retina | 768 px | no, canvas baked in | 60 KB | visible ringing along cable edges |
| Product hero | 1536 px | no, canvas baked in | 140 KB | banding across a matte plastic shell |
| Composable cutout | source width | yes, alpha preserved | none, cold tier | fringing when recomposited |
Baking the canvas into every card-facing derivative is the decision that buys the bandwidth: an opaque image drops the alpha channel entirely, and the encoder stops spending bits describing transparency that the page will never show. The cost is that those files are only valid for one canvas color, which is fine once they're disposable and cheap to regenerate. The one place alpha survives is the composable cutout that lives next to the original.
Here's the compositing and budget search, condensed to the part that matters:
from dataclasses import dataclass
from io import BytesIO
from PIL import Image
@dataclass(frozen=True)
class Slot:
width: int
budget_kb: int
fmt: str
LADDER = (
Slot(384, 24, "WEBP"),
Slot(768, 60, "WEBP"),
Slot(1536, 140, "WEBP"),
)
def compose(original_path: str, mask_path: str, canvas: tuple[int, int, int]) -> Image.Image:
"""Product over today's canvas color. The canvas is an argument, never baked into a master."""
original = Image.open(original_path).convert("RGB")
mask = Image.open(mask_path).convert("L")
if mask.size != original.size:
mask = mask.resize(original.size, Image.LANCZOS)
plate = Image.new("RGB", original.size, canvas)
return Image.composite(original, plate, mask) # mask=255 keeps the product
def encode(img: Image.Image, slot: Slot) -> bytes:
height = round(img.height * slot.width / img.width)
resized = img.resize((slot.width, height), Image.LANCZOS)
for quality in (82, 74, 66, 58, 50):
buf = BytesIO()
resized.save(buf, format=slot.fmt, quality=quality, method=6)
data = buf.getvalue()
if len(data) <= slot.budget_kb * 1024:
return data
raise BudgetMiss(f"{slot.width}px stayed above {slot.budget_kb} KB down to quality 50")
The descent stops at the first quality that fits, so a plain black probe lands at 82 and a busy keyboard with 87 legends lands wherever it has to. The BudgetMiss at the bottom is deliberate. A pipeline that silently drops to quality 30 to satisfy a budget will ship a blurred hero and nobody notices for a quarter; an exception routes that SKU to a human instead. When it fires, the fix is usually art direction rather than encoder settings: crop tighter, or reshoot the product larger in frame.
Reprocessing the whole catalog without babysitting it
Derivative keys have to be deterministic, or a backfill is unresumable. Hash the inputs that actually change the output bytes:
import hashlib
ENCODER_REV = "cutout-v3"
def derivative_key(original_digest: str, mask_digest: str, slot: Slot, canvas_hex: str) -> str:
spec = f"{original_digest}|{mask_digest}|{slot.width}|{slot.fmt}|{canvas_hex}|{ENCODER_REV}"
return f"derived/{hashlib.sha256(spec.encode()).hexdigest()[:32]}.{slot.fmt.lower()}"
Bump the canvas color or the encoder revision and every key moves, which means the new generation writes alongside the old one instead of over it. Nothing is invalidated mid-flight, no purge storm, no half-refreshed grid for users holding a warm cache. You render, you verify, you flip the manifest, and a lifecycle rule reaps the previous generation a week later. Re-running an interrupted backfill is free, because keys already written are keys you can skip.
Two operational habits carry most of the weight here. First, the job writes one row per derivative — slot, byte size, chosen quality, encoder revision — so the byte distribution per slot is queryable instead of anecdotal, and a regression after an encoder bump shows up as the 95th percentile drifting past its budget. Second, a golden set of about 40 SKUs, deliberately loaded with the hard cases — braided cables, black-on-black enclosures, a mesh speaker grille — gets recomposited on every pipeline change and diffed against the approved masks. It's the same instinct as an eval harness for a model: a small, nasty, fixed sample beats a large average.
What to measure before copying this ladder
Measure three things in your own catalog: the byte size distribution per slot against the budget you claim to have, how often the budget search bottoms out, and the full cost of one reprocess split into inference and compositing. That third number decides whether keeping masks is worth the storage, and it's the number most teams have never computed.
The catch is that this architecture is built for churn. If your products arrive already shot on a uniform sweep, your layout has been stable for years, and there's one card size, a single rendered file per SKU is the correct answer and everything above is overhead you'll maintain for nothing. Manifest, mask store, budget search, golden set — all of it earns its keep only when the ladder moves. It's also a poor fit for user-generated photos you have no rights to retain indefinitely; there, a retention window on originals matters more than reprocessing convenience, and a deletion path that provably removes both original and derivatives is the harder engineering problem. Video is its own discipline too: stick with a streaming toolchain built for it rather than stretching a still-image ladder across formats it was never designed for.
Both libvips and ImageMagick expose the compositing and alpha primitives this needs, with libvips streaming large images in strips to keep memory flat per job while ImageMagick's in-memory model is heavier. That difference stops being academic the moment a backfill runs 16 workers on one box.
Keep the original. Keep the mask. Everything else is a cache.
Further reading
- MDN — Image file type and format guide: https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
- WHATWG HTML — srcset and sizes attributes: https://html.spec.whatwg.org/multipage/images.html#srcset-attributes
- RFC 9111 — HTTP Caching: https://www.rfc-editor.org/rfc/rfc9111.html
- Pillow — image file formats (WebP options): https://pillow.readthedocs.io/en/stable/handbook/image-file-formats.html
- libvips — memory and streaming model: https://www.libvips.org/API/current/How-it-opens-files.html
- WebP container specification (alpha channel): https://developers.google.com/speed/webp/docs/riff_container
- AV1 Image File Format (AVIF) specification: https://aomediacodec.github.io/av1-avif/
Top comments (0)