DEV Community

KasimirBerg5341
KasimirBerg5341

Posted on

5 Checks for Transformation-Not-Found Deploys — Environment Names and Missing Setup

Short answer: treat a transformation-not-found error after deploy as configuration drift until proven otherwise: publish one immutable transformation manifest before application traffic moves, validate the exact environment-qualified names against it, and keep generated image keys independent of the deployment environment.

For a B2B marketplace that smart-crops listing images into 1:1, 4:3, and 16:9 variants, this is an architecture decision, not a retry problem. A retry can repeat the same bad lookup while creating extra cache misses; a manifest check identifies whether the deployed application and the media control plane agree.

The storage rule is blunt: preserve one original, derive only the aspect ratios that a live placement requests, and make every derivative key deterministic. Don't let staging and production names leak into that key.

Names are contracts.

1. How should you debug a transformation not found error after deploy?

Start at the boundary where the application turns a business intent such as listing-card-square into a stored transformation identifier. Capture four values from the same request: deployment environment, requested logical name, resolved transformation name, and source object key. A request ID is useful, but those four values answer the actual question.

Then compare the resolved name with the manifest shipped for that environment. If the application asks for prod-listing-card-square while the manifest contains production-listing-card-square, the mismatch is already explained; inspecting pixels, codecs, or crop coordinates would waste time. If the name exists, move one boundary outward and verify that the application is reading the intended manifest revision, then one boundary inward and verify that the incoming logical name is permitted. This order is deliberate because it tests cheap, deterministic state before invoking an image operation.

Keep the error classification narrow. unknown_logical_name means code requested a transformation outside the contract. environment_mapping_missing means deployment configuration could not resolve a logical name. manifest_entry_missing means resolution succeeded but the published manifest lacks that entry. source_missing is a storage problem and should not be mislabeled as a transformation problem. The user-facing response can remain generic while logs and metrics retain this bounded reason code.

Stop there first.

I'm not sure which boundary is wrong in a system without its deploy revision and resolved name; no amount of general advice can replace those two observations. What is knowable is the investigation order, and it should not begin by purging the cache.

2. Name the invariants and failure boundaries

The decision record needs invariants that survive deployments. The logical transformation names are application contracts. The environment mapping is deployment configuration. The manifest is published control-plane state. Generated objects are data-plane state. Combining those layers into one string may feel convenient, but it makes a rename look like new image content and can strand perfectly valid derivatives behind obsolete cache keys.

For the marketplace example, a useful invariant is: the tuple (source_version, logical_transform, transform_revision, output_format) identifies a derivative. Environment is absent. Two deployments that use the same transformation revision should address the same object, while a real crop-policy change increments transform_revision and produces a new object. This makes rollback predictable and prevents a blue/green deployment from doubling derivative storage merely because one side calls the environment prod and the other calls it production.

There are limits. Content-derived keys require a stable source version, and lazy generation means the first request for a new listing/ratio pair pays the transformation latency. Pre-generating everything avoids that first-request path but is not suitable when most uploaded listings never appear in every placement; it spends compute and storage on cold derivatives. If every asset is guaranteed to be shown in all three fixed ratios, pre-generation can be the simpler and more observable choice. Your mileage may vary with listing churn and cache retention, so measure requested ratio cardinality rather than guessing.

The failure boundary should also preserve the original. Image format support differs across browsers and file types, and MDN's format guide documents those compatibility considerations. A derivative policy therefore should choose an output format from a declared client capability, while retaining the source object as the durable input for later reprocessing. A missing transformation definition must never trigger an overwrite of that source.

3. Compare setup strategies before changing code

Strategy Deploy behavior Storage and cache effect Main limitation
Runtime-created definitions The application creates or updates names while serving traffic A rename can split cache keys unless identifiers are normalized Startup and request handling now share control-plane responsibility
Environment-owned mutable names Each environment maintains human-readable definitions independently Easy to create duplicate derivatives when names drift Equality of names does not prove equality of crop policy
Versioned manifest published before traffic CI publishes an immutable mapping, then the application validates its revision Stable logical keys preserve cache reuse; policy revisions invalidate intentionally Requires a deploy gate and manifest retention for rollback
Fully local crop rules The application carries crop parameters and performs transformation itself Storage keys can be deterministic, but compute and cache operations belong to the team Operational burden is justified only when local control matters

The default choice here is a versioned manifest published before traffic. It separates setup from request serving, supports an explicit rollback target, and gives deployment automation a finite assertion: every required logical name exists at the expected revision. It doesn't prove that a crop is aesthetically good. That belongs in fixture-based visual review, because a valid 1:1 crop can still remove the product from a marketplace thumbnail. A deploy gate should use a small fixture set that represents the failure modes the product actually has: a portrait listing photo, a landscape photo, an image with the subject near an edge, and an image whose orientation metadata changes the displayed direction. The gate checks definition presence and key determinism; a separate visual check reviews crop quality. Mixing those tests produces vague failures that are harder to route to the owning team.

Cost follows from cardinality. For N source versions, R requested ratios, F negotiated output formats, and V active crop revisions, the upper bound on derivative identities is N × R × F × V; the cache may hold fewer, but a naming error can add an accidental environment dimension. That extra dimension is the one this design removes. No invented savings percentage is needed to justify it.

4. Put the critical path in executable policy

The request path should resolve names and build keys without mutating transformation setup. This Python example is intentionally provider-independent: the manifest has already been published, and the function either returns a deterministic plan or a bounded configuration error.

from dataclasses import dataclass
from hashlib import sha256
from typing import Mapping


class MediaConfigurationError(Exception):
    pass


@dataclass(frozen=True)
class TransformSpec:
    revision: str
    width: int
    height: int


RATIOS: Mapping[str, TransformSpec] = {
    "listing-square": TransformSpec(revision="crop-v3", width=1200, height=1200),
    "listing-standard": TransformSpec(revision="crop-v3", width=1200, height=900),
    "listing-wide": TransformSpec(revision="crop-v3", width=1600, height=900),
}

ENVIRONMENT_NAMES: Mapping[str, Mapping[str, str]] = {
    "staging": {name: f"staging-{name}" for name in RATIOS},
    "production": {name: f"production-{name}" for name in RATIOS},
}


def plan_derivative(
    environment: str,
    logical_name: str,
    source_key: str,
    source_version: str,
    output_format: str,
    published_manifest: Mapping[str, str],
) -> dict[str, object]:
    environment_map = ENVIRONMENT_NAMES.get(environment)
    if environment_map is None:
        raise MediaConfigurationError("environment_mapping_missing")

    spec = RATIOS.get(logical_name)
    resolved_name = environment_map.get(logical_name)
    if spec is None or resolved_name is None:
        raise MediaConfigurationError("unknown_logical_name")

    if published_manifest.get(resolved_name) != spec.revision:
        raise MediaConfigurationError("manifest_entry_missing")

    identity = "|".join(
        [source_key, source_version, logical_name, spec.revision, output_format]
    )
    derivative_id = sha256(identity.encode("utf-8")).hexdigest()

    return {
        "resolved_name": resolved_name,
        "revision": spec.revision,
        "width": spec.width,
        "height": spec.height,
        "object_key": f"derived/{derivative_id}.{output_format}",
    }
Enter fullscreen mode Exit fullscreen mode

Notice what the key excludes: the resolved environment name. That name is still logged because it explains setup drift, but it cannot fragment stored derivatives. Also notice that the example validates a revision rather than mere membership. Two environments can contain the same logical entry while pointing at different crop parameters; presence alone would bless a silent policy mismatch.

Cache keys aren't.

Production code should emit one counter per bounded error class and attach the deployment revision, manifest revision, and logical name as structured fields. Avoid putting raw listing identifiers into low-cardinality metric labels; keep them in sampled logs keyed by request ID. An alert on manifest_entry_missing immediately after traffic shifts is actionable, while an undifferentiated transformation error rate is just a symptom.

5. Record the rejected option and its valid use case

This ADR rejects request-time creation of transformation definitions for the marketplace path. The catch is ownership: a read request should not need permission to mutate shared setup, and concurrent application versions should not race to reinterpret a mutable name. Retries don't repair a disagreement about configuration. They repeat it.

Request-time creation is still valid for a single-tenant internal tool where one process owns both definition lifecycle and rendering, traffic is low, and a cache split has little consequence. Stick with fully local transformation rules when regulatory or data-residency constraints require all image processing inside infrastructure your team operates, or when the crop algorithm itself is proprietary enough to justify the maintenance load. Those cases trade a larger operational surface for control, which can be the correct exchange.

For the multi-tenant listing system, deployment ordering is the cleaner contract: publish manifest, verify required names and revisions, shift traffic, and retain the prior manifest for rollback. After the shift, compare error counters by deployment revision and watch derivative-key cardinality. A sudden new environment-shaped prefix is evidence that the key contract regressed even if images still render.

The final decision is vendor-neutral: separate logical intent from environment lookup, version the transformation policy, validate setup before traffic, and address derivatives by content-relevant inputs. That resolves the missing setup step while keeping storage and cache behavior explainable.

References

Top comments (0)