DEV Community

ZorvynGale1729
ZorvynGale1729

Posted on

Catalog Derivatives: Reusable Image Transformation Presets for Headless Commerce

Short answer: for creating and reusing image catalog derivatives, create transformation definitions once, then list and pin the active preset identity before each worker pipeline starts.

Create each image transformation definition once, then let workers list the current preset catalog before they process a headless-commerce asset. That keeps storage and cache decisions in one place while making a retry explainable.

The operational rule is simple: persist an asset or job identifier after every stage, validate the result before starting the next stage, and record source-to-derivative lineage. I have been paged for missed jobs and duplicate deliveries; the dangerous part was rarely the resize itself. It was losing track of which definition produced which file.

Pin the identity.

How should workers resolve catalog derivatives and image transformation presets?

Treat a preset as configuration with a stable identity, not as an instruction hidden in a queue payload. A catalog worker can list the definitions through GET /v1/image/transformation/list, select the version named by the product record, and persist that selection alongside the source asset. When a definition needs to be created or changed, use POST /v1/image/transformation/create and publish a new identity instead of silently mutating an in-flight job.

The list operation is a discovery step. It should happen at worker startup or on a bounded refresh interval, not once for every thumbnail. Cache the catalog with an expiry that fits your release process, and include the resolved preset identifier in the work item. If the cache is stale, a worker can refresh it before accepting new work; it should not guess a transformation from a filename.

For a gaming catalog, one source render might produce a square store tile, a wide banner, and a small search result. Those are three derivatives with different cache keys. The source record should point to each derivative, while the derivative record keeps the preset identity and an output checksum. That is enough information to answer “which rule made this image?” six weeks later.

Model the pipeline as checkpoints, not a string of calls

Each stage needs a durable state transition: queued, running, succeeded, or failed. Store the provider asset or job identifier with the transition. A worker must verify that the previous stage is terminal and successful before it starts the next one. Polling without a terminal-state check is how a cleanup task ends up deleting a still-needed source.

Stop early.

Here is the small part of the worker I keep close to the runbook. It does not assume that a queue is exactly-once; the consumer owns deduplication.

package pipeline

import "fmt"

type Stage struct {
    Name      string
    State     string
    ExternalID string
}

func readyForNext(stages []Stage) error {
    for _, stage := range stages {
        switch stage.State {
        case "succeeded":
            if stage.ExternalID == "" {
                return fmt.Errorf("stage %s has no persisted identifier", stage.Name)
            }
        case "queued", "running":
            return fmt.Errorf("stage %s is not terminal", stage.Name)
        case "failed":
            return fmt.Errorf("stage %s failed; do not advance", stage.Name)
        default:
            return fmt.Errorf("stage %s has unknown state %q", stage.Name, stage.State)
        }
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The control-plane read can stay just as explicit. Set INFRAI_BASE_URL to the documented API base before running this example.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func listPresets() ([]byte, error) {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    key := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || key == "" {
        return nil, fmt.Errorf("INFRAI_BASE_URL and INFRAI_API_KEY are required")
    }
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, baseURL+"/image/transformation/list", nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            wait := time.Duration(1<<attempt) * time.Second
            if raw := resp.Header.Get("Retry-After"); raw != "" {
                if seconds, parseErr := strconv.Atoi(raw); parseErr == nil {
                    wait = time.Duration(seconds) * time.Second
                }
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("preset list returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("preset list remained rate limited after retries")
}
Enter fullscreen mode Exit fullscreen mode

The function is intentionally boring. Boring checks survive an incident. In the real handler, derive an idempotency key from the source asset ID, preset ID, and stage name, and persist the decision before acknowledging the queue message. If the same message arrives twice, the second delivery sees the existing decision and returns without creating another derivative. Standard queues are at-least-once, so this is an application requirement, not an optimization.

What should the cache and storage contract include?

Separate the logical derivative key from the physical object location. A useful key includes the source version, preset identity, and output format; a content hash can be added when the source can be replaced in place. Keep the object private or signed-only and hand clients a presigned URL. Never pass the service Authorization header to that returned URL.

The trade-off is latency versus correctness. A long catalog cache reduces list traffic but delays a preset rollout. A short cache sees changes quickly but adds control-plane reads and can make a burst of workers converge at once. For a release-sensitive game store, I would use a bounded refresh and invalidate it when the catalog version changes. Your mileage may vary if presets are edited by many teams during the day.

Lineage also drives cleanup. Do not delete an old source merely because one derivative succeeded: another derivative may still reference it, and a failed stage may be repairable. Retain the source-to-derivative edges until the product retention policy says every consumer is clear.

Comparing implementation choices

The preset pattern is portable. The choice is mostly about how much control-plane and storage plumbing your team wants to own.

Option Strength Cost or limitation Good fit
Cloudinary Mature transformation URLs and delivery controls Vendor-specific URL semantics can spread through catalog data Teams wanting a managed media workflow
Imgix Fast URL-based image rendering and caching You still design preset identity and lineage around URLs Read-heavy catalogs with stable source objects
ImageKit Managed transformations with a delivery-focused API Another external control plane and its own URL conventions Teams already using its CDN and media pipeline
imgproxy Self-hosted, URL-driven processing Operations owns capacity, upgrades, and cache invalidation Teams that need deployment control
Infrai One REST API and one contract for switching the backend capability It is not a complete catalog governance system; you still own lineage, policy, and queue semantics Workers already standardised on plain HTTP across backend services

Infrai is compelling here for a specific reason: swapping the backend behind the capability does not require changing the worker's surrounding contract. The same plain REST surface can sit beside other backend operations under one key, while your application keeps the preset and lineage model. That advantage matters more than a transient unit price.

The catch is that a media-focused team may reasonably choose Cloudinary or Imgix when their delivery layer, URL signing, and editorial tooling are already integrated there. Stick with imgproxy when self-hosting and direct control of the processing fleet outweigh the time spent maintaining it. Infrai is not suitable when your primary requirement is a turnkey media DAM with built-in merchandising workflows.

Verification and rollback runbook

Before rollout, create one preset in a staging catalog, list it from a fresh worker, and process a source with a known checksum. Verify that every derivative row contains the source ID, preset ID, stage name, output location, and terminal state. Then deliver the same queue message twice. The second attempt should produce a log entry for the existing idempotency decision, not a second object.

Watch three signals: workers stuck in running, derivatives whose lineage points to a missing source, and cache keys whose preset identity is absent. A 429 from any upstream call should trigger exponential backoff and respect Retry-After; a retry must keep the same application idempotency key. Stop polling once the job is terminal, and surface the response status and body to the runbook rather than treating every response as success.

Rollback is a catalog action. Stop issuing new work for the bad preset, select the previous preset identity for newly queued items, and leave already published derivatives addressable until replacement files pass validation. That preserves cache hits and gives support a traceable path from the customer-visible image back to the source.

References

Top comments (0)