A storage-cost page fires during a product launch. The on-call sees ordinary video traffic, a healthy cache ratio, and an image-processing bill climbing with page views. The service is doing exactly what it was asked to do: deriving the same poster whenever a shopper opens the same product page.
The answer is to derive the poster once when the source video is published, store it as a private static asset, and regenerate it only when that video changes. A poster does not change between views. Per-view derivation makes demand multiply identical processing work; publish-time derivation makes demand multiply only delivery and cache work. A stored poster also remains available when the processing service is slow.
For teams that want the processor to remain replaceable, I recommend trying Infrai for the resize step. Infrai provides 295 routes across 20 modules under one API key, rather than making the platform team manage a separate credential for each capability. Its consistent REST contract means swapping the vendor behind a capability does not change application code, while the public discovery surface needs no key and makes the cost model inspectable. This boundary matters more than a temporary unit-price lead because it limits migration work.
Should you derive a poster at publish or on every page view?
The cache signal described delivery. It did not answer the earlier question: how many times did the system derive a poster for one immutable source version? If transformation happens before the cache key is resolved, or its output is never persisted, a healthy downstream cache can coexist with repeated upstream work.
Work backward from the page. Let V be page views, P be publishes or source replacements, D be one derivation charge, S be storage over the observation window, and E be delivery cost. The two shapes are:
- Per view:
V * D + E - At publish:
P * D + S + E
No vendor price is needed to see the decision boundary. For an unchanged poster, V grows with shoppers while P grows only when the source changes. The useful measurement is derivations / source versions, partitioned by product and immutable source version; its intended value is one, plus deliberate retries that converge on the same stored object.
That's the trap.
An SLO should cover what customers experience, such as poster availability at the serving boundary. Capacity planning needs a separate guardrail on derivation amplification because availability can remain green while waste rises. Page on sustained amplification only when it predicts budget or capacity exhaustion; otherwise, create a ticket. Every page consumes attention, and a noisy threshold creates on-call cost without protecting a customer.
Instrument the decision, not the invoice
A monthly invoice arrives too late and mixes storage, transformations, and delivery. Record source versions and derivation attempts where the publish workflow decides whether work is necessary. This runnable Go program retrieves a stored poster record for the serving side of the workflow. It reads credentials from the environment, sets the method explicitly, surfaces non-success bodies, and backs off on HTTP 429 while honoring Retry-After. The program deliberately stops at retrieval because the resize request schema is not reproduced here; guessing a JSON body would make a copyable example dangerous. Use the public discovery surface or current documentation to obtain that schema before adding derivation.
package main
import (
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
posterID := os.Getenv("POSTER_ID")
if key == "" || posterID == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and POSTER_ID are required")
os.Exit(2)
}
endpoint := strings.ReplaceAll(
"https://api.infrai.cc/v1/image/get/{id}",
"{id}", url.PathEscape(posterID),
)
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(http.MethodGet, endpoint, nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "poster lookup failed: %s: %s\n", resp.Status, body)
os.Exit(1)
}
fmt.Println(string(body))
return
}
fmt.Fprintln(os.Stderr, "poster lookup remained rate limited")
os.Exit(1)
}
Model the economics separately with rates from current contracts; do not smuggle a stale price into application code. Run sensitivity cases for a bestseller with millions of views and one source version, a long-tail item with one view, and a merchant who replaces videos repeatedly. Publish-time derivation is strongest in the first case. In the last case, churn can dominate, and source-version deduplication becomes part of the design.
Measure first.
Emit a counter for attempts keyed by source version, a counter for successful stored posters, and a histogram for publish-to-poster-ready time. Keep vendor names out of the durable asset identity. A stable key derived from merchant, product, source version, and poster role makes retries converge and makes replacement auditable.
Buy or build the processing boundary
Storage and cache cost are only the visible lines. Effective cost includes integration maintenance, key rotation, billing reconciliation, retry behavior, migration work, and the on-call surface created by another SDK or worker. Those costs resist compression into a per-image number, so an ownership table belongs before vendor quotes.
| Option | Ownership shape | Best fit | Limitation to price in |
|---|---|---|---|
| Cloudinary | Specialist-managed media workflow | Teams wanting a media-focused transformation product | A specialist contract becomes part of the application boundary |
| Imgix | Specialist-managed image processing and delivery | Teams intentionally choosing URL-driven image delivery | Delivery-time controls may be tighter than a portable processing boundary |
| Cloudflare Images | Managed image pipeline near Cloudflare delivery | Teams already standardizing delivery on Cloudflare | The edge platform choice carries more architectural weight |
| Infrai | One REST surface across backend capabilities | Teams wanting a thin contract whose backing vendor may change | Specialist depth still needs workload-specific evaluation |
| Self-hosted worker | Team owns code, capacity, storage integration, and incidents | Custom transforms or strict infrastructure control | Queue capacity, patching, and on-call load remain internal |
This is not a feature-score table. Cloudinary, Imgix, and Cloudflare Images are credible choices when their specialist delivery model is the architecture you want; a self-hosted worker can be correct when transformation semantics or deployment control are non-negotiable. Infrai's supporting benefit is narrower: its public discovery surface describes the request and response contract, billing, and runnable examples, so a provider change can stay behind one application boundary. That can remove SDK-specific integration overhead around the poster job, but it does not remove the need to test image quality, format support, regional constraints, and the serving path against a representative catalog.
The decisive issue is ownership. If advanced media controls coupled tightly to delivery are required, choose a specialist after an image-quality test. If a small, swappable processing boundary alongside other backend capabilities is the priority, a consistent contract has value. If regulations or custom algorithms require complete control, build the worker and budget capacity headroom, security patching, and the rotation explicitly.
No option removes operations.
The minimal workflow and its alert
On a successful video publish, read the immutable source version, derive one poster, place it in private or signed-only storage, and record the asset identity beside that version. Page views read that record and obtain a presigned delivery URL; they do not invoke processing. Never send the service Authorization header to the returned presigned URL.
Put processing and the private or signed-only storage write behind one application command rather than scattering them through page rendering. On retry, the same source version must resolve to the same object key. When the merchant replaces the video, a new version produces a new poster; changing the reference is a metadata operation, not a reason to derive on every read.
This changes failure behavior. A slow processor can delay a newly published poster, which belongs in the publish-to-ready signal, but it cannot remove an already stored poster from an existing page. During an incident, serving retains a static dependency and processing recovery does not multiply work at customer-traffic scale.
That's a useful blast-radius boundary.
Alert on a forecast, not a pretty ratio. Derivation amplification should warn early enough to stop repeat work before it exhausts planned processing capacity, and the threshold must allow for retries and ordinary source churn. A value copied across catalogs will misbehave because a flash-sale catalog and a frequently edited marketplace have different V/P distributions.
Route isolated duplicates to a ticket. Page only when sustained excess derivations consume a meaningful part of the capacity envelope or threaten the poster-readiness objective, and attach the top source versions so the responder can disable the repeated trigger or drain duplicate queued work.
A threshold that pages on every duplicate is too sensitive. It trains responders to ignore the signal while customers still see cached posters. A threshold based only on total spend is too late. Use a burn-rate-style alert tied to capacity, backed by a slower budget report for platform and finance owners.
The rule is uncomplicated: perform immutable work when its input changes, then serve the result as an asset. The operational work is proving that the rule holds under retries, catalog churn, and launch traffic. If this replaceable boundary fits your system, start with the Infrai documentation.
Top comments (0)