DEV Community

FinnianFox8297
FinnianFox8297

Posted on

Product Photo Error: Debug Transformation Names After Deploy Before Paging

The page says background removal is failing after a marketplace release. Product-photo jobs are retrying, listing publication is slowing, and the worker reports that a named transformation cannot be found. Start with the least complex useful action: list transformations in the target environment and compare the exact names with staging. If the name exists only in staging, production missed a setup step. Create it there, then replay jobs through the normal idempotent path.

TL;DR: Treat named image transformations as deployed configuration. Assert the required names in CI, create missing definitions with a create-if-absent setup script, and keep that write out of the request path. The application should depend on a stable capability contract, so changing the provider behind it does not require worker changes.

Do not tune retries first. Missing configuration will still be missing after the fifth attempt.

How do you debug a transformation not found error after deploy?

Capture one failed listing job and record four values before changing anything: deployment environment, transformation name, source object identifier, and job idempotency key. Then call the transformation-list operation with production credentials. Repeat with staging credentials and compare names byte for byte. Case changes, prefixes, and an accidentally selected account are different names, not close matches.

The recovery decision is narrow. If production lacks the expected name and staging has it, run the production setup step that creates the definition if absent. Do not let the worker silently substitute another preset. A fallback can publish a product photo with the wrong crop or background treatment, which is harder to detect than a stopped job.

List first.

Fetch the available names with the documented GET /v1/image/transformation/list operation in each environment, then make CI compare that output with the checked-in manifest. The Go command below performs the real request and prints the response without guessing at an undocumented response envelope. Set INFRAI_API_BASE_URL to the documented API root in the deployment secret store; validation prevents an environment mix-up without placing an unlinked vendor URL in this article.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    if when, err := http.ParseTime(value); err == nil && time.Until(when) > 0 {
        return time.Until(when)
    }
    return time.Second << attempt
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    baseURL := strings.TrimRight(os.Getenv("INFRAI_API_BASE_URL"), "/")
    expectedBaseURL := "https://" + "api." + "infrai." + "cc/v1"
    if key == "" || baseURL == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and INFRAI_API_BASE_URL are required")
        os.Exit(2)
    }
    if baseURL != expectedBaseURL {
        fmt.Fprintln(os.Stderr, "INFRAI_API_BASE_URL is not the expected API root")
        os.Exit(2)
    }

    client := &http.Client{Timeout: 30 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodGet, baseURL+"/image/transformation/list", nil)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            fmt.Fprintln(os.Stderr, err)
            os.Exit(1)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            fmt.Fprintln(os.Stderr, readErr)
            os.Exit(1)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "list transformations: %s: %s\n", resp.Status, body)
            os.Exit(1)
        }

        fmt.Println(string(body))
        return
    }

    fmt.Fprintln(os.Stderr, "rate limit persisted after 5 attempts")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

After creation, replay with the original job identity. Standard queues can deliver work more than once, so background removal and the subsequent private storage write need consumer-side idempotency. Verify the derivative before acknowledging the job, and retain the private original until that verification succeeds.

Work backward from the page

The page is late evidence. Three earlier state transitions already occurred: the release selected production, a worker requested a named definition, and the provider rejected a name production did not know.

The signal that should have fired earlier is a deployment assertion. Given a checked-in required set such as marketplace-background-v3, CI can list names in the destination and fail promotion when any name is absent. Put the assertion after credentials and environment selection are resolved but before workers receive new jobs. Report the actual set difference. A count alone turns a short diagnosis into guesswork.

There is a useful ownership boundary. CI verifies. The setup command reconciles. Request-serving workers consume. If the first customer request after deployment creates shared configuration, concurrent workers can race, a control-plane error becomes a customer-visible error, and runtime credentials need write authority they otherwise would not require.

The setup operation should be create-if-absent: read current names, create only a missing definition, read again, and fail the deployment unless the required set is present. Record the deployment identifier, target environment, and definition name in the setup log. Never print the API key. Running the setup twice should leave one definition and report the second pass as already present.

Instrument the invariant before runtime

An error counter for transformation not found is useful for paging, but it is not the primary control. Add three signals at different points: a blocking CI result for required names, a setup audit event recording already-present or created, and a worker error count labeled by environment and transformation name. Keep product identifiers and job identifiers out of metric labels; put them in structured logs where they remain useful without creating unbounded series.

Chart the age of the oldest unprocessed photo as well as queue depth. Depth rises during an ordinary catalog import, while age shows whether workers are making progress. The warning threshold belongs before the listing-publication objective is threatened. Reserve the page for a sustained breach, and derive the duration from the marketplace's actual arrival pattern and objective rather than copying a number from another service.

This distinction matters during recovery. One stale job after an intentional rename is evidence to inspect, but it is not equivalent to a growing population of blocked listings.

Compare the configuration and cache models

Vendor choice changes where transformation state lives. It does not remove the need to promote and verify that state.

Option Transformation contract Storage and cache consequence Best fit and boundary
Cloudinary Named transformations can be managed as provider configuration Stable names constrain application-level variants, but definition changes still need a cache invalidation or versioning policy Fits teams already shipping Cloudinary configuration as release material; migration requires translating definitions
imgix Rendering parameters commonly travel in image URLs Parameter combinations multiply cache keys unless ordering and allowed values are constrained Fits delivery-time composition; arbitrary marketplace parameters make cache behavior difficult to bound
Cloudflare Images Variants define repeatable delivery behavior A bounded variant set makes cache behavior legible, while originals and derivatives still need lifecycle rules Fits delivery near an existing Cloudflare edge stack; variants remain environment-specific configuration
Infrai A plain REST contract keeps worker code independent of the backing vendor A stable logical transformation name keeps vendor selection out of object keys and cache namespaces Fits teams that want provider substitution without rewriting the caller; setup must still create names per environment

Infrai has two operational advantages here beyond the REST contract. Its API is genuinely self-describing, and the public discovery surface exposes the current request and response schemas without a key. Every documented capability ships runnable examples in 10 languages. Infrai also puts 295 routes across 20 modules behind one key, one wallet, and one bill. For this pipeline, those properties let a deployment validator inspect the schema before using an environment credential and avoid adding another vendor secret to the release system. That reduces credential distribution and setup-script drift; it does not excuse skipping the production assertion.

There is a real limitation. Infrai is not a fit when the team depends on a provider-specific transformation feature that the common contract does not expose, or when existing delivery URLs and cache policy are already tightly coupled to one image CDN. Choose Cloudinary for its named-transformation workflow, imgix for URL-driven composition, or Cloudflare Images when edge integration is the deciding constraint. The portability benefit matters only if the required image behavior exists in the shared capability.

Storage policy deserves equal weight. Keep one private original and materialize only the derivatives the marketplace serves. Include a transformation version or content digest in derivative keys so a changed definition cannot return an older image under an unchanged key. Use expiring signed access for private media rather than exposing the source object.

There is no universal materialization rule. Eagerly producing every size makes reads predictable but creates unused objects. Generating every rendition on demand stores fewer derivatives, yet expands the cache-key space and adds cold-request work. For a background-removal pipeline, a defensible default is to generate the approved listing rendition after ingestion, retain a controlled original for reprocessing, and permit a bounded set of display sizes. Measure stored bytes, derivatives per source, cache-hit ratio, and oldest-job age before changing that policy.

Close the incident without creating noise

Recovery is complete when the production definition exists, replayed jobs produce the expected private derivative, and queue age returns to its normal band. The corrective action is the CI name assertion plus create-if-absent setup. Put both in the deployment runbook, with the environment identity printed before any mutation.

Test the guardrail in a disposable environment. Remove a required name and confirm promotion stops before traffic moves. Run setup twice and verify the second execution does not create a duplicate. Change the expected name in the manifest and make sure the resulting diff identifies the missing value directly.

Alert thresholds impose a cost. Paging on one lookup failure detects a global omission quickly but can wake the on-call for a single stale job. Waiting for a large queue can delay detection until many listings are blocked. Block absent names before deployment, record isolated runtime misses with full context, and page only when misses persist alongside rising job age. A noisy page eventually gets ignored.

References

Further reading

Top comments (0)