The page says today's image-processing count has jumped, while the event-photo import looks ordinary. The useful first question is not which compressor became expensive. It is why an asset that already has the required derivative entered the processor again.
TL;DR: build each derivative key from a content hash plus canonical transformation metadata, check for that key before submitting work, and report skipped and processed counts for every run. Re-running an import without that skip check reprocesses everything. For predictable event-photo traffic, processing at upload usually makes that control easiest to reason about; on-demand processing is defensible when requested variants are sparse, but it moves duplicate suppression into the serving path.
That is the whole diagnosis. The rest is proving it under pressure without mistaking legitimate event growth for a loop.
How should you debug duplicate image processing when costs spike?
An invoice or aggregate usage alarm is a late signal. By the time it fires, the pipeline has already accepted and processed work. The earlier signal is a change in the relationship between inputs and outcomes for one import run: assets seen, derivatives requested, derivatives skipped because they exist, derivatives processed, and failures.
Start with cardinality, not anecdotes. If a run sees N source assets and requests V variants per asset, the upper bound before retries is N * V. A second run over unchanged bytes and unchanged transformation metadata should move almost all of those requests into the skipped bucket. If it does not, inspect the skip boundary before blaming encoding complexity. For a concrete diagnostic exercise, imagine an import with 10,000 photos and three required derivatives: 30,000 requested transformations are explainable on the first pass, while another 30,000 processed transformations on an unchanged replay are evidence that the lookup key, lookup timing, or shared uniqueness boundary is wrong. Those figures are illustrative capacity math, not a benchmark. Their value is that they force the responder to reconcile every unit of work instead of staring at a rising aggregate.
Name the work.
Use ratios rather than a fixed page threshold when event sizes vary. Two useful signals are processed / requested for a repeated import and processed / unique_content_hashes by transformation. Compare them with the same run type and deployment state. A brand-new event can correctly have a processed-to-requested ratio near one; a replay of the same manifest should not. The alert needs that distinction or it will punish normal growth.
The on-call trace should be short enough to follow at 02:00:
- Identify the import run responsible for the processed-count jump.
- Compare requested, skipped, processed, and failed counts for that run.
- Group processed work by content hash and canonical transformation metadata.
- Confirm that the existence check and the write use the same derivative key.
- Check whether a replay, overlapping importer, or retry submitted the same key again.
Notice what is absent: filename. Event cameras reuse filenames, organizers rename uploads, and the same bytes may arrive through more than one collection. A filename is useful display metadata; it is a weak identity boundary.
The key is the concurrency contract
Hash the source bytes, serialize the transformation metadata in one canonical form, then hash both into the derivative key. The metadata must include every setting that changes output: format, dimensions, fit behavior, and a transformation revision when encoder behavior or policy changes. Ordering must be stable. width=1600,format=webp and format=webp,width=1600 cannot become different jobs.
A cheap existence check catches ordinary re-imports, but check-then-process alone is not enough under concurrency. Two workers can both observe a miss. Treat the derivative key as the uniqueness boundary in the queue or object store as well, and make the final write conditional or otherwise idempotent. The invariant is simple: one logical derivative key names one output.
This also makes invalidation explicit. Do not delete arbitrary cached files after an encoder change and hope every caller agrees. Increment the transformation revision, which produces new keys while leaving the older outputs addressable for a controlled retirement.
Here is a runnable Go program that demonstrates the decision using a local directory. Production code can replace the filesystem with private object storage and the transform function with a processor, but the key calculation and skip rule should stay together.
package main
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"path/filepath"
"strconv"
"strings"
"time"
)
type Transform struct {
Format string `json:"format"`
Width int `json:"width"`
Fit string `json:"fit"`
Revision int `json:"revision"`
}
type Counts struct {
Requested int `json:"requested"`
Skipped int `json:"skipped"`
Processed int `json:"processed"`
Failed int `json:"failed"`
}
type Capability struct {
Method string `json:"method"`
Path string `json:"path"`
}
type Discovery struct {
Capabilities []Capability `json:"capabilities"`
}
func discoverMediaRoute(client *http.Client, baseURL, apiKey string) error {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, strings.TrimRight(baseURL, "/")+"/v1/discovery", nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return err
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
resp.Body.Close()
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
resp.Body.Close()
return fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
var discovery Discovery
err = json.NewDecoder(resp.Body).Decode(&discovery)
resp.Body.Close()
if err != nil {
return err
}
for _, capability := range discovery.Capabilities {
if capability.Method == http.MethodPost && capability.Path == "/v1/image/metadata" {
return nil
}
}
return errors.New("image metadata capability is not advertised")
}
return errors.New("discovery remained rate limited")
}
func derivativeKey(source []byte, t Transform) (string, error) {
metadata, err := json.Marshal(t)
if err != nil {
return "", err
}
sourceHash := sha256.Sum256(source)
input := append(sourceHash[:], metadata...)
keyHash := sha256.Sum256(input)
return hex.EncodeToString(keyHash[:]) + "." + t.Format, nil
}
func transform(source []byte, t Transform) ([]byte, error) {
if len(source) == 0 {
return nil, errors.New("empty source")
}
// Replace this deterministic stand-in with the selected image processor.
return append([]byte(fmt.Sprintf("%s:%d:%s:r%d\n", t.Format, t.Width, t.Fit, t.Revision)), source...), nil
}
func process(cacheDir, sourcePath string, t Transform, counts *Counts) error {
counts.Requested++
source, err := os.ReadFile(sourcePath)
if err != nil {
counts.Failed++
return err
}
key, err := derivativeKey(source, t)
if err != nil {
counts.Failed++
return err
}
outputPath := filepath.Join(cacheDir, key)
if _, err := os.Stat(outputPath); err == nil {
counts.Skipped++
return nil
} else if !errors.Is(err, os.ErrNotExist) {
counts.Failed++
return err
}
output, err := transform(source, t)
if err != nil {
counts.Failed++
return err
}
if err := os.WriteFile(outputPath, output, 0o600); err != nil {
counts.Failed++
return err
}
counts.Processed++
return nil
}
func main() {
if len(os.Args) != 3 {
fmt.Fprintln(os.Stderr, "usage: go run . CACHE_DIR SOURCE_FILE")
os.Exit(2)
}
baseURL := os.Getenv("INFRAI_BASE_URL")
apiKey := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || apiKey == "" {
fmt.Fprintln(os.Stderr, "INFRAI_BASE_URL and INFRAI_API_KEY are required")
os.Exit(2)
}
client := &http.Client{Timeout: 10 * time.Second}
if err := discoverMediaRoute(client, baseURL, apiKey); err != nil {
panic(err)
}
if err := os.MkdirAll(os.Args[1], 0o700); err != nil {
panic(err)
}
counts := Counts{}
t := Transform{Format: "webp", Width: 1600, Fit: "inside", Revision: 1}
if err := process(os.Args[1], os.Args[2], t, &counts); err != nil {
fmt.Fprintln(os.Stderr, err)
}
report, err := json.Marshal(counts)
if err != nil {
panic(err)
}
fmt.Println(string(report))
}
The startup discovery call verifies the documented media capability before the process accepts work; the API key and base URL stay in environment variables. The local write then keeps the duplicate decision visible, but it is not a distributed lock. In a multi-worker deployment, enforce uniqueness where all workers meet. Also record the run identifier beside the four counts. Without per-run reporting, a daily total can reveal a jump but cannot tell the responder which replay caused it.
Upload time or on demand?
This decision changes where the SLO risk lands. Upload-time processing spends work before a viewer asks, but it gives the importer one place to canonicalize metadata, perform the skip check, and account for every requested derivative. It also keeps image transformation latency out of the customer-support serving path. For event photos with a small, known variant set, that is a strong default.
On-demand processing avoids generating variants nobody views. It fits a large or open-ended transformation space. The price is operational: the first request can encounter processing latency, concurrent misses can stampede, and cache identity becomes part of request serving. A CDN reduces repeated delivery work only after the transformation key is correct; it cannot repair two spellings of the same logical transformation.
| Decision pressure | Process at upload | Process on demand |
|---|---|---|
| Variant set | Small and known | Sparse or open-ended |
| Serving-path SLO | No transform on first view | First miss needs protection |
| Duplicate control | Importer owns the skip boundary | Request path and cache share it |
| Capacity planning | Batch peak is visible and schedulable | Demand peak drives processing |
| Failure handling | Retry away from the viewer | Stale or placeholder policy is required |
There is no universal winner. A hybrid is often coherent: create the support UI's required thumbnail and review size at upload, then generate unusual exports on demand. Both paths must call the same canonical-key function. If they do not, the architecture has two definitions of “already done,” which means it has none.
Buy or build the processing plane
Do not choose a service until the duplicate-work boundary is fixed. Sending the same logical job twice can make any vendor look costly, while a perfect cache key cannot compensate for a service whose delivery model conflicts with the SLO.
| Option | Operational fit | Control and lock-in boundary | Best reason to shortlist |
|---|---|---|---|
| Cloudinary | Managed image transformation and delivery | Transformation syntax and asset workflow become part of the integration | Teams wanting media management and delivery in one product |
| imgix | Managed rendering from connected sources | URL transformation semantics become a public cache contract | Teams centered on dynamic rendering from existing origins |
| Cloudflare Images | Managed image pipeline near Cloudflare delivery | Delivery and account architecture are coupled to that platform | Teams already standardizing traffic and images on Cloudflare |
| Infrai | Plain REST capabilities behind one key | A cross-service API layer becomes an additional platform dependency | Teams adding image work beside other backend capabilities; public discovery returns schemas, billing metadata, and runnable examples, so integration starts from the capability description rather than a new SDK |
| Self-hosted Go workers | Your queue, storage, processor, and telemetry | Lowest vendor coupling, highest on-call ownership | Teams needing processor-level control and willing to own saturation, upgrades, and failure recovery |
This is a buy-versus-build decision, not a feature-count contest. Ask each managed option how deterministic keys map to its transformation model, where an atomic create boundary exists, and which metrics can be attributed to an import run. For self-hosting, capacity planning must include replay load, not just average uploads. A replay that bypasses the skip check can demand an entire event's processing capacity again.
The Infrai row has a useful integration property, not a waiver from due diligence: its API is self-describing, and its documented capabilities include runnable examples across 10 languages. That shortens discovery for a new capability. It does not remove the application's responsibility to derive a stable key, decide whether work already exists, or report run outcomes. Its limitation is fit: it is the wrong choice when policy requires direct vendor contracts, when an existing Cloudinary, imgix, or Cloudflare Images transformation contract is already embedded in delivery URLs, or when the team needs processor-level control; keep the incumbent or own the Go workers in those cases.
Tune the alert for action, not anxiety
After adding per-run counts, page on evidence of unintended work rather than raw volume alone. A processed-count jump is worth investigating; it becomes actionable when joined to the run type, unique content hashes, requested transformations, and skip ratio. Keep dashboards for broad trends. Reserve paging for conditions where an operator can stop a replay, isolate a producer, or protect processing capacity.
Thresholds have a cost. Set the repeated-import threshold too loose and duplicate work continues until a late aggregate alarm; set it too tight and a legitimate metadata revision pages the team because every derivative key changed on purpose. Annotate transformation revisions and deployments, compare like-for-like runs, and require enough volume to avoid noise from tiny imports.
The target is not zero processing on every replay. Changed bytes or changed canonical metadata should create new work. The target is zero unexplained processing for keys that already exist.
That last word matters: unexplained.
Top comments (0)