Pick two identifiers before you pick a watermark library: one for the source asset you ingested, one for every marked preview derived from it. The source stays byte-identical forever, and the marked preview becomes a cache entry with a TTL and a rebuild recipe. Replacing the source with a watermarked copy costs nothing on the day you do it and everything on the day someone needs the clean original — for a dispute file, a re-crop, a legal hold — because that's the one file you can't regenerate.
The mark itself is cheap. The copies are what you pay for.
Where the money goes in a preview pipeline
The system worth thinking through is a merchant catalog inside a payments platform. Sellers upload product photos during onboarding, a matting step strips the background so every tile sits on the same neutral card, and the cleaned cutout goes back to the seller for approval before buyers ever see it. Those approval previews leave the building watermarked, because they travel through notification email, support tickets and the occasional screenshot in a group chat before anyone signs off. I spend most of my time on email and OTP delivery, so that detail is not incidental to me: the moment an image is attached to a message, you have lost control of where it ends up, and a preview that leaks unmarked is a merchant's product shot circulating with your platform's implied blessing on it.
Now count the objects. Forty thousand SKUs, three aspect ratios, two delivery formats, marked and unmarked variants of each: 480,000 derived objects against 40,000 masters. Storage is the boring half of that bill. The interesting half is cache behaviour — a catalog page that paints forty tiles at once will either hit the CDN forty times or generate forty cutouts on demand, and the difference between those two outcomes is the difference between a flat monthly line item and a spiky one that tracks your crawl traffic.
That fan-out is the constraint. The watermarking question is downstream of it.
What does it cost to watermark previews without replacing the source assets?
Roughly one extra object per shape you actually serve, plus a rebuild you can trigger on demand. Not one extra copy of everything — that distinction is the whole design. If a rendition is derived deterministically from (source id, recipe, parameters), then the derived object is disposable: you can evict it, let it expire, or purge a whole prefix after a recipe change, and the next request rebuilds it. The master never participates in any of that.
| Approach | Storage cost | Cache behaviour | Blast radius when the recipe changes |
|---|---|---|---|
| Mark burned into the ingested file | One copy per asset | Trivial, one key | Unrecoverable: the clean original is gone |
| Derive on write, precompute every shape | N copies per asset, paid up front | High hit ratio, no cold start | Rebuild and purge N objects per asset |
| Derive on read, cache with a TTL | Only what is requested | Cold first hit per key | Expire the prefix, traffic rebuilds what matters |
Most catalogs I have reasoned about land on a split rather than a pure strategy: precompute the two renditions that appear on every listing page, derive everything else lazily. The tail of a product catalog is long and cold, and precomputing marked previews for SKUs nobody has opened in eight months is the clearest waste in the whole pipeline.
Protection cost has a second component that never shows up in the storage line. Every marked rendition you cache is a rendition your CDN must not confuse with the unmarked one. Get that wrong and you have not overspent — you have leaked.
A derivation chain you can replay
The rule that keeps this honest: derive the storage key from the recipe, so a recipe change can't silently reuse an object rendered under the old one. Same discipline as an idempotency key on a payment or an OTP send — one deterministic id per intended effect, never a mutation in place.
import hashlib
import io
from dataclasses import dataclass
from PIL import Image
RECIPE_VERSION = "v3" # bump when the mark, the crop or the encoder settings change
@dataclass(frozen=True)
class Rendition:
source_id: str
width: int
fmt: str # "webp" or "jpeg"
marked: bool
def derivative_key(r: Rendition) -> str:
raw = f"{RECIPE_VERSION}|{r.source_id}|{r.width}|{r.fmt}|{int(r.marked)}"
return f"derived/{hashlib.sha256(raw.encode()).hexdigest()[:24]}.{r.fmt}"
def render(store, r: Rendition) -> str:
key = derivative_key(r)
if store.exists(key):
return key
with store.open(r.source_id) as f: # the ingested master, opened read-only
base = Image.open(f).convert("RGBA")
cutout = remove_background(base) # matting stage; the source is not written back
cutout.thumbnail((r.width, r.width))
if r.marked:
cutout.alpha_composite(mark_layer(cutout.size), (0, 0))
buf = io.BytesIO()
cutout.convert("RGB").save(buf, format=r.fmt.upper(), quality=72)
store.put_if_absent(key, buf.getvalue(),
cache_control="public, max-age=31536000, immutable")
return key
Two details in there earn their place. put_if_absent means two workers racing on the same cold key write the same bytes and neither corrupts the other, which matters more than it sounds when a bot crawls your catalog and forty tiles miss at once. And the long max-age is only safe because the key contains the recipe hash — an immutable object under a content-addressed name is fine, an immutable object under /preview/{sku}.webp is a year-long mistake.
The public path deserves its own guard rather than a comment:
def public_url(store, r: Rendition, ttl: int = 300) -> str:
if not r.marked:
raise PermissionError("unmarked renditions are not served on the public path")
return store.signed_url(derivative_key(r), ttl=ttl)
Tooling for the marking stage is mostly a memory-profile decision. libvips streams through an image in chunks and keeps peak memory low per worker, which is what you want when a queue consumer might pick up a 6000-pixel master. ImageMagick is easier to script and has a composite operation that most teams already know, at the cost of holding fuller rasters in memory. If the catalog later accepts short clips, the equivalent stage is an ffmpeg overlay filter, and the derivation-key idea carries over unchanged.
Cache keys are where the protection actually leaks
Everything above is straightforward until content negotiation shows up. If your CDN serves AVIF to browsers that advertise it and WebP to everyone else from the same URL, the response must carry Vary: Accept, or a shared cache is entitled to hand the stored variant to a client that asked for something different. RFC 9111 spells out how a stored response is matched against a new request; the failure mode is not theoretical, and it is much easier to avoid by putting the format in the key — as the snippet above does — than by trusting every intermediary in the path to honour a header correctly.
Three edge cases have burned people I have worked near:
- A signed URL with a 5-minute TTL cached by an intermediary for an hour. The signature expires, the cached copy doesn't, and now you have an object whose access control is a lie.
- Negative caching of a 404 during a backfill, so the tile stays broken for the length of the TTL after the rendition finally lands.
- A purge that clears the marked prefix but leaves an unmarked debug rendition someone added for QA six months earlier.
That last one is the compliance problem in miniature. Watermarking is a control, and a control that is enforced in the rendering code but not in the storage layout will eventually be routed around by a colleague who needs a clean file for a deck. Put the unmarked derivatives in a separate bucket or prefix with its own policy, keep the master under object lock or an equivalent write-once rule for as long as your dispute and audit windows require — in card disputes that is measured in months, not days — and make the public path physically unable to reach either.
On provenance: embedding signed metadata alongside the visible mark is a reasonable second layer, and C2PA is the specification most of that work has converged on. It survives a screenshot exactly as well as any metadata does, which is to say not at all. Invisible robust watermarking is the other direction, and I wouldn't promise you a number for how well any particular scheme survives a crop-and-recompress round trip — published robustness results vary widely by algorithm and by how motivated the attacker is. If the threat model is a determined adversary rather than casual reuse, that gap is worth measuring on your own assets before you rely on it.
Rolling it out, and when to do the opposite
Backfilling 480,000 objects on a Tuesday isn't a plan. Ship the derivation chain behind a dual-read: try the derived key, fall back to whatever legacy path exists, and log which one answered. Once the hit ratio on the new path is where you expect, precompute the hot renditions — in practice the top few percent of SKUs by view count covers most of the traffic — and let the tail stay lazy. Then flip the public path to reject unmarked renditions, and only after that turn on the lifecycle rule that expires derived objects. Expire early and you will spend the savings on rebuild compute.
The catch is that this design assumes rebuilds are cheap and previews are disposable. Both assumptions break in real places. A pipeline whose matting step runs a heavyweight segmentation model can make a cold miss expensive enough that derive-on-read stops being a saving at all, and then precomputing everything is the honest answer. If your renditions are effectively permanent — regulated archives, printed catalog masters, anything with a retention obligation attached to the delivered file — a cache with a TTL is not a good fit, and you should be storing those as first-class objects with their own lifecycle. Very small catalogs don't need any of this; a few hundred assets marked once at upload, with the originals kept in a separate write-once bucket, gets the same protection with a fraction of the machinery.
And if a single delivered file must carry the mark permanently, burn it in — into a derived copy, still never into what you ingested.
Top comments (0)