Derive a fixed ladder of variants at upload time, and use request-time resizing only for the sizes you genuinely cannot predict. The cost that decides this is not CPU. It's cache: a resize-on-request URL lets clients invent sizes, every invented size is a new cache key, and an unbounded variant count means your origin keeps re-rendering the same seller photo at 401px because somebody typed the wrong number into a template.
| Approach | Variant count | What the cache holds | Typical tooling | Where it hurts |
|---|---|---|---|---|
| Upload-time ladder | Fixed — you pick 3–5 widths | A small warm set per image | libvips or sharp in your own worker, or a REST resize call | Adding a width is a reprocessing job over the originals |
| Request-time, open URL params | Unbounded | A long tail of one-hit keys | imgix, Cloudinary, ImageKit URL APIs | Hit rate decays, origin re-renders, egress climbs |
| Request-time, signed allowlist | Bounded by the allowlist | Same warm set, plus late arrivals | Cloudflare Images variants, signed imgix URLs | You maintain signing code and a size registry |
| Hybrid | Fixed core, small tail | Warm core, cold exceptions | Ladder in a worker, CDN transforms for oddities | Two code paths that both have to stay honest |
For a marketplace library that also gets auto-tagged for search, I'd ship the hybrid and let the ladder carry the layouts you actually render. The exception path exists, it's signed, and it's audited. Everything else comes from the fixed set.
Two consumers, one derivative set. That framing is what makes this comparison different from the usual thumbnail debate.
Where the bandwidth goes when a model reads your images too
A listing photo in a marketplace has two audiences. Buyers see it three times — a 256px grid cell, a 768px card, a 1600px detail view — and each of those is a derivative you can name in advance. The auto-tagger sees it once, and what it sees decides how the listing gets found.
That second consumer is the one people forget when they budget bandwidth.
Tag quality tracks input resolution up to a point, and the useful tags in a marketplace live in fine detail: the brand stamped on a derailleur, the model number printed on a box, the fabric weave that separates "linen" from "cotton blend" in a search facet. A 256px thumbnail gets you "bicycle" and "outdoors" — technically correct tags that nobody searches for. Feed the original 4000×3000 JPEG instead and you get the detail back, but now every image in the catalog is a multi-megabyte read from storage into the tagger, once at ingest and again every time you change the model. Somewhere in between is a tier that keeps the tags you care about and cuts the bytes by an order of magnitude. Finding it is a measurement, not an opinion, and I'll come back to how I'd run it.
Pick that tier on purpose. If the tagger just grabs "whatever URL the CDN happened to have warm," your tag quality becomes a function of cache state, which is a wonderfully hard bug to reproduce six months later.
Should you resize on upload or on request when every variant gets cached?
Upload-time derivation gives you a fixed cost and a fixed set of variants. Request-time resizing gives you flexibility and an unbounded variant count. That's the whole trade, and the cache is where it shows up.
Count the keys. A ladder of four widths in two formats is eight objects per image — a set you can enumerate, warm, and invalidate. Open the URL parameters instead and the key space is the product of width, height, quality, fit mode and format, which is thousands of reachable combinations per image even before anyone gets creative. Clients will invent sizes if the URLs let them. A designer ships w=401 in one template, a partner's srcset generator emits sixteen widths instead of your four, and each one is a fresh key, a fresh origin render, and a fresh full-resolution read of the original.
None of those renders are wrong. They're just work you didn't need to do, paid for in cold-cache latency on the exact pages you most want to be fast.
The quieter advantage of upload-time derivation is that it's reprocessable. Keep the original private and immutable, and a new tier is a batch job over storage rather than a client-side migration — add a 384px width for a redesigned grid, backfill, ship. I'd rather run that batch than negotiate URL formats with three frontend teams. Request-time earns its place when the size honestly cannot be predicted: partner embeds where you don't control the layout, art-directed crops chosen per listing, or an API where third parties pick their own dimensions.
Quality per byte: sizing the tier your tagger reads
Here's the test I'd run before arguing about architecture. Take a few hundred listings sampled from your worst categories — not the studio-lit ones, the phone photos shot in a garage. Tag each one at 256px, 512px and 1024px, and again from the original. Then compare the tag sets against the original's, and take the smallest tier where agreement stops improving.
That plateau is your tier. Where it lands depends on your catalog, so I'm not going to pretend there's a universal answer — a furniture marketplace and a trading-card marketplace will not land in the same place, and your mileage may vary even between categories.
Then do the arithmetic that actually sets the bill: tier bytes × images ingested per day × re-tagging passes per year. The re-tagging multiplier is the one that bites. Every model upgrade is a full re-read of the library, and if your tagger reads originals, that pass costs the same as the ingest did.
One upload, one ladder, in TypeScript
If you don't want to run your own resize worker, the request-time products above will do it at the edge, and a plain API call will do it at ingest. Infrai's image resize route is one HTTP POST with no SDK to install — GET /v1/discovery/image.resize hands back the JSON Schema for the request, the response shape, and a runnable example, so wiring it up is reading one endpoint rather than learning a client library. The same Infrai key also covers the storage and queue calls wrapped around that resize, so the ingest path carries one credential instead of three.
The catch is that it doesn't expose signed transform URLs at the edge the way the CDN-native products do. If your plan is transform-on-request at the CDN, stick with imgix or Cloudflare Images and don't fight it.
type Tier = { width: number; label: string };
const LADDER: Tier[] = [
{ width: 256, label: "grid" },
{ width: 768, label: "card" },
{ width: 1600, label: "detail" },
];
const TAGGER_TIER = 768;
const BASE = requireEnv("INFRAI_BASE_URL");
const KEY = requireEnv("INFRAI_API_KEY");
function requireEnv(name: string): string {
const value = process.env[name];
if (!value) throw new Error(`missing env var: ${name}`);
return value;
}
const sleep = (ms: number) => new Promise((done) => setTimeout(done, ms));
async function resize(source: string, width: number, idempotencyKey: string): Promise<unknown> {
for (let attempt = 0; ; attempt++) {
const res = await fetch(`${BASE}/image/resize`, {
method: "POST",
headers: {
authorization: `Bearer ${KEY}`,
"content-type": "application/json",
"idempotency-key": idempotencyKey,
},
body: JSON.stringify({ image: source, width, fit: "inside", format: "webp", store: true }),
});
if (res.status === 429 && attempt < 4) {
const retryAfter = Number(res.headers.get("retry-after") ?? 0);
await sleep(retryAfter > 0 ? retryAfter * 1000 : 500 * 2 ** attempt);
continue;
}
const text = await res.text();
if (!res.ok) throw new Error(`resize ${res.status}: ${text.slice(0, 200)}`);
return JSON.parse(text);
}
}
// sha256 is the content hash of the original, so a replayed upload event
// resolves to the same derivative instead of creating a second one.
export async function deriveLadder(source: string, sha256: string) {
const variants: Record<string, unknown> = {};
for (const tier of LADDER) {
variants[tier.label] = await resize(source, tier.width, `${sha256}:w${tier.width}`);
}
return { variants, taggerInput: variants[LADDER.find((t) => t.width === TAGGER_TIER)!.label] };
}
Three things in there are not decoration. The idempotency key is derived from the content hash and the width, so a redelivered upload webhook can't double-charge you for the same derivative. The 429 branch honours Retry-After before falling back to exponential backoff — ingest bursts are spiky by nature, and a tight retry loop turns a queue drain into an outage of your own making. And the tagger reads a named tier from the ladder, not an ad-hoc URL, which is the part that keeps tag quality reproducible.
When request-time resizing is the right call
There's one case where the fixed ladder is the wrong default, and it isn't the one people usually cite. If most of your library is never viewed — a long-tail marketplace where old listings sit untouched — then deriving four variants per upload means paying to render images nobody will request. Derive lazily, cache aggressively, and accept the cold-start cost on first view. The tagging tier still gets derived at ingest, because that one is guaranteed to be read.
The other case is genuine unpredictability: third-party consumers, embeddable widgets, or a partner API where the dimension is theirs to choose. Bound it with a signed allowlist anyway. An allowlist of twenty sizes is still a bounded key space; an open w parameter is not.
Everything else — your own grid, your own cards, your own detail pages, your own tagger — is predictable, and predictable variants belong in a ladder you build once at upload. Fixed set, warm cache, reprocessable originals. That's the boring answer, and it's the one I'd defend in review.
Top comments (0)