DEV Community

AlgernonCross4103
AlgernonCross4103

Posted on

Python Healthtech Duplicate Image Derivatives (When Processing Costs Surge)

Short answer: put an atomic, content-addressed skip check immediately before image transformation, then publish a derivative only after its bytes and metadata are complete. In a healthtech event-photo service, this makes retries converge on one stored result instead of paying to compress the same asset again under different job IDs.

The deciding constraint is storage and cache cost. A queue can deliver the same logical request twice, an upload callback can race a reconciliation scan, and two filenames can contain identical bytes. None of those paths should create another derivative. Treat that as an architecture invariant, not as a lucky property of the worker.

How do I debug duplicate image processing costs?

Start at the bill, but do not stop there. Separate source storage, derivative storage, transformation executions, cache fills, and egress in the cost dashboard. If transformation count rises faster than the number of distinct derivative keys, the pipeline is repeating work. If stored bytes rise instead, key construction or cleanup is suspect. A lower cache-hit ratio with stable origin counts points toward unstable URLs or metadata, not necessarily duplicate compression.

The useful unit is a derivative intent: source content plus every setting that can change the output. For an event photo, that may include the source digest, normalized orientation, width, output media type, quality policy version, and encoder version. A request ID is not part of the image. Neither is a retry number.

This distinction catches an expensive edge case. Two attendees may upload the same photo under stage.jpg and IMG_8041.JPG; filename deduplication misses it. Conversely, one filename can be replaced with different bytes, so filename deduplication can incorrectly serve stale clinical-event imagery. Compare content identity, not labels.

Names lie.

One caution belongs near the front: do not place patient identifiers, event names, or other sensitive context in a public object key. A digest-based key is opaque, but authorization and retention controls still apply to the source and every derivative.

Invariants and failure boundaries

The decision record has four invariants:

  1. Equal source bytes and equal transformation parameters resolve to one derivative key.
  2. A cacheable object is visible only after its complete payload is stored.
  3. A retry can observe ready, wait on processing, or safely take over expired work; it cannot blindly start another transform.
  4. Failures remain retryable without publishing a partial derivative as success.

The race sits between “I checked” and “I started.” A plain lookup followed by an insert is not a skip check when two workers run concurrently. Both can see absence. The claim must therefore be an atomic conditional write in the system of record, backed by a uniqueness constraint on the derivative key.

Keep the database row and object payload as separate failure boundaries. The row coordinates ownership; object storage holds bytes. Write to a temporary object name, validate the encoded result, promote or copy it to the deterministic final key, and only then mark the row ready. A worker that dies before readiness leaves recoverable state rather than a convincing half-success.

Short leases are tempting, especially during an incident. They are also dangerous when a large source takes longer than expected. Record a lease expiry, renew it during active work, and measure takeovers. The correct lease duration comes from observed high-percentile transform time plus operational margin, not from a tidy round number.

The trade-off is real: coordination adds a database write and state to repair.

One comparison, centered on cost behavior

Strategy Duplicate boundary Concurrency behavior Storage/cache consequence Best fit
Request-ID key One delivery attempt Every retry can win Repeated objects and cold cache entries Temporary diagnostics only
Source filename plus preset One mutable label Races unless separately locked Stale collisions and missed cross-name duplicates Controlled, immutable imports
Content digest plus versioned parameters One derivative intent Converges when atomically claimed Stable object and cache keys Retry-heavy asynchronous pipelines
Perceptual similarity Visually similar inputs Requires a threshold and candidate search May merge images that are not byte-equivalent Review tooling, not exact idempotency

I choose the third option for this ADR. It costs one source digest calculation and a coordination write, but it aligns the key with the reusable artifact. It also makes a cache miss intelligible: either this exact intent has never completed, or the object and registry disagree. This design is not suitable for a latency-sensitive synchronous path that cannot afford hashing the full source before work begins; a trusted upstream digest or a different ingestion boundary is needed there. It is also a poor fit when transformations intentionally include nondeterministic data, because identical intent may not produce reusable bytes unless that input becomes part of the key.

No hash prevents a race by itself. The database uniqueness rule is what turns two simultaneous claims into one owner and one observer.

That is the hinge.

The critical path in Python

The following code keeps vendor details behind small interfaces. The repository method must implement claim as an atomic insert-if-absent operation. The object store must not expose the final key until put_complete has finished.

from dataclasses import dataclass
from hashlib import sha256
import json
from typing import Protocol


@dataclass(frozen=True)
class Transform:
    width: int
    media_type: str
    quality_policy: str
    encoder_version: str


class Registry(Protocol):
    def claim(self, key: str) -> str:
        """Atomically return 'owner', 'processing', or 'ready'."""

    def mark_ready(self, key: str, byte_count: int) -> None: ...
    def mark_failed(self, key: str) -> None: ...


class ObjectStore(Protocol):
    def put_complete(self, key: str, payload: bytes, media_type: str) -> None: ...


def derivative_key(source: bytes, transform: Transform) -> str:
    source_digest = sha256(source).hexdigest()
    parameters = json.dumps(
        {
            "encoder_version": transform.encoder_version,
            "media_type": transform.media_type,
            "quality_policy": transform.quality_policy,
            "width": transform.width,
        },
        sort_keys=True,
        separators=(",", ":"),
    ).encode("utf-8")
    intent_digest = sha256(source_digest.encode("ascii") + b"\x00" + parameters).hexdigest()
    extension = {"image/jpeg": "jpg", "image/png": "png", "image/webp": "webp"}[
        transform.media_type
    ]
    return f"derivatives/{intent_digest[:2]}/{intent_digest}.{extension}"


def process_once(source: bytes, transform: Transform, registry: Registry,
                 objects: ObjectStore, encode) -> str:
    key = derivative_key(source, transform)
    state = registry.claim(key)

    if state == "ready":
        return key
    if state == "processing":
        raise RuntimeError("derivative is already being produced; retry with backoff")

    try:
        payload = encode(source, transform)
        if not payload:
            raise ValueError("encoder returned an empty derivative")
        objects.put_complete(key, payload, transform.media_type)
        registry.mark_ready(key, len(payload))
        return key
    except Exception:
        registry.mark_failed(key)
        raise
Enter fullscreen mode Exit fullscreen mode

There is deliberate friction in the parameter object. If a quality policy changes, its version changes too. Otherwise an old compressed image may be returned under a key that now claims to represent a new policy. The same rule applies to encoder upgrades when they can alter output bytes.

Media type belongs in the intent because JPEG, PNG, and WebP have different characteristics and browser support histories. Select a type from actual content requirements and supported clients, then keep the extension and served Content-Type consistent. The MDN image-format guide is a useful compatibility reference; it is not a substitute for testing the particular photographs, transparency needs, metadata policy, and clients in this service.

Debugging the spike without creating another one

Instrument transitions, not raw filenames. Each worker should emit the derivative-key prefix, claim outcome, source byte count, output byte count, transform duration, policy version, and terminal state. Keep sensitive upload metadata out of labels because high-cardinality, identifying telemetry is both costly and hard to govern.

Three ratios are enough to localize most duplicate-work incidents: claims per distinct intent, encodes per successful intent, and ready records per final object. They should be computed over the same time window and policy version. A claims spike with stable encodes means retries are being skipped correctly. An encode spike means the atomic claim, lease, or key normalization is failing. Extra final objects suggest nondeterministic names or publication outside the guarded path.

Use a controlled replay before changing the key scheme. Start 2 workers with the same bytes under two names, deliver the same job concurrently, retry after an injected encoder failure, and vary exactly one transform parameter. The expected result is 1 ready object for the first three successful paths and a separate object for the changed parameter. Then test a worker death after upload but before mark_ready; reconciliation should verify the final object before deciding whether to adopt it or retry. This small matrix is more informative than replaying a large production batch because each variation tests one boundary: identity, concurrency, recovery, or intentional divergence.

Do not “fix” the graph by caching exceptions indefinitely. A failed encode is not a completed derivative. Record enough failure state to apply bounded backoff and investigate poison inputs, while preserving a path to retry after the cause is corrected.

Rejected option and the narrow case where it works

I reject a preflight object-exists check as the primary guard. It has a time-of-check/time-of-use gap, cannot distinguish a complete object from poorly published state without extra metadata, and encourages every worker to hit storage before coordination. Under concurrent delivery, two workers can both miss and both encode.

It still has a valid use case: reconciliation. A periodic repair process can compare ready rows with final objects, identify abandoned temporary uploads, and examine expired processing leases. In a single-threaded backfill where the input manifest is immutable and no other producer writes the namespace, an existence check may also be an acceptable operational shortcut. Those conditions should be explicit and temporary.

The production rule stays simple: derive identity from bytes and versioned intent, claim it atomically, and publish once. This contains transformation spend, keeps cache keys stable, and gives incident responders evidence about which boundary failed instead of another pile of differently named copies.

References

Top comments (0)