DEV Community

RaffertyBarrett4726
RaffertyBarrett4726

Posted on

Reusable Product Assets with Go: Reliable Background Removal Under Upload Bursts

The page fires after a course-merchandising upload: the original garment photo exists, but one or more fashion catalog cutouts still show the studio wall because background removal missed its publication deadline. The worker queue looks busy, so the tempting response is to add workers. That can create duplicate assets without telling the on-call whether quality, scheduling, or delivery actually failed. The alert needs to identify the oldest affected item, its last completed stage, and the batch deadline while there is still time to recover.

Short answer: keep one lossless transparent master per fashion item, generate delivery variants asynchronously, and alert on the age of the oldest incomplete asset rather than raw queue depth. Gate completion on both edge quality and variant delivery. This is the least complex design that keeps background removal reusable while making upload bursts operable.

The central trade-off is quality versus bandwidth. Don't make the removal model, the queue, and browser encoding one opaque step. A clean master preserves the expensive decision about which pixels belong to the product; smaller derivatives can then change as the catalog layout and browser mix change.

Queue depth is not enough.

How should Go schedule background removal for reusable fashion catalog assets?

Treat an upload as a small state machine, not a callback that happens to write three files. The useful states are received, cutout_ready, variants_ready, and published. A job may run more than once. Publication may not.

The idempotency key should identify the source content and transformation policy, not a queue message ID. In practical terms, hash the original bytes plus a versioned policy containing the removal model, crop rule, and quality profile. Retries then converge on the same object names. A revised crop policy gets a new key and leaves the previous catalog asset intact until promotion.

Here is the core scheduling boundary in Go. The repository and queue are deliberately generic; the important behavior is the compare-and-set claim and deterministic job key.

package cutouts

import (
    "context"
    "crypto/sha256"
    "encoding/hex"
    "errors"
)

var ErrAlreadyScheduled = errors.New("asset already scheduled")

type AssetStore interface {
    Claim(ctx context.Context, key string) (bool, error)
}

type Queue interface {
    Enqueue(ctx context.Context, job Job) error
}

type Job struct {
    Key       string
    SourceURI string
    Policy    string
}

func Schedule(ctx context.Context, store AssetStore, queue Queue, source []byte, sourceURI, policy string) (string, error) {
    h := sha256.New()
    _, _ = h.Write(source)
    _, _ = h.Write([]byte("\x00" + policy))
    key := hex.EncodeToString(h.Sum(nil))

    claimed, err := store.Claim(ctx, key)
    if err != nil {
        return "", err
    }
    if !claimed {
        return key, ErrAlreadyScheduled
    }
    if err := queue.Enqueue(ctx, Job{Key: key, SourceURI: sourceURI, Policy: policy}); err != nil {
        return "", err
    }
    return key, nil
}
Enter fullscreen mode Exit fullscreen mode

There is a catch: a permanent claim made before enqueue needs reconciliation if the process exits between those operations. Use a transactional outbox when the claim and job record share a database; otherwise give claims a lease and run a sweeper that republishes expired, incomplete work. The second option accepts more duplicate delivery attempts, so every output write and publish transition still has to be idempotent.

I've been paged by missed jobs and duplicate deliveries. The lesson is plain: “the queue accepted it” is not a customer outcome.

For background removal, mark cutout_ready only after checking the artifact. A practical automated gate can verify dimensions, a nonempty alpha channel, and a plausible foreground area. Those checks catch empty or fully opaque outputs, but they can't prove that lace, hair, shadows, and pale fabric edges look right. Keep a reviewed fixture set spanning those cases, run it whenever the transformation policy changes, and require a human visual check for ambiguous high-value products. I'm not sure one foreground-area range can serve every fashion category; measured distributions from your own approved assets are what should set that bound.

Work backward from the page

The first alert should answer three questions without a dashboard tour: which publication batch is late, how old its oldest incomplete item is, and which stage owns the delay. Queue length alone answers none of them. Ten large source images can represent more work than a thousand cached retries, and a quiet queue can coexist with catalog items stuck between storage and publication.

Work backward from the visible failure. The page fired because the catalog had crossed its promised publication deadline. Before that, the oldest incomplete item was aging. Before that, one stage's completion rate had fallen below its arrival rate or its retries had stopped making progress. Follow one asset through the timeline: upload acceptance creates durable state, scheduling claims the policy-specific key, removal writes the transparent master, encoding creates the requested renditions, validation marks those renditions complete, and one manifest promotion makes them visible together. A missing timestamp narrows the search to the transition before it; a timestamp with no later progress points at the next stage. This trace also distinguishes an old retry from genuinely old customer work, which raw queue age often conflates. Instrument every transition, then put the customer-facing deadline at the top of the alert.

Use a monotonic per-asset lifecycle where possible, and emit a counter for every attempted transition with from, to, and a bounded reason. Record queue wait and processing duration separately. The labels should describe stages and policies, not asset IDs; IDs belong in structured logs and traces, where they don't create unbounded metric cardinality.

An illustrative service-level objective might require 99% of a publication batch to reach published within 180 seconds of upload acceptance. Those numbers are an example, not a benchmark. Set them from the editor workflow, measured source sizes, and the actual catalog launch deadline. The alert can then combine a burn-rate signal with an oldest-item guard, while a lower-severity ticket catches a single stranded asset.

This is the instrumentation change that usually matters: emit the age of incomplete work from durable state, not from a worker's memory.

package cutouts

import (
    "context"
    "time"
)

type IncompleteReader interface {
    OldestAcceptedAt(ctx context.Context, stage string) (time.Time, bool, error)
}

type Gauge interface {
    Set(stage string, seconds float64)
}

func ObserveOldestAge(ctx context.Context, now time.Time, stages []string, reader IncompleteReader, gauge Gauge) error {
    for _, stage := range stages {
        acceptedAt, found, err := reader.OldestAcceptedAt(ctx, stage)
        if err != nil {
            return err
        }
        age := 0.0
        if found {
            age = now.Sub(acceptedAt).Seconds()
            if age < 0 {
                age = 0
            }
        }
        gauge.Set(stage, age)
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The runbook starts with the oldest asset key and its last successful transition. Check whether new items are progressing through that stage, compare arrival and completion rates, and inspect retry reasons. Scale workers only when wait time is rising and processing capacity is the constraint. If processing time rose after a policy deployment, stop promoting the new policy and let the prior version drain; if publication alone is late, adding removal workers just increases pressure on the wrong stage.

Preserve quality once, spend bandwidth by context

The reusable artifact is the transparent master with enough detail for future crops and sizes. It is not necessarily the file a browser should download. Delivery variants should be derived from that master for the display box, density, and acceptable visual loss, with dimensions encoded in the object identity so a retry cannot overwrite a different rendition.

MDN's media format guide documents that image formats differ in compression behavior, transparency support, and browser compatibility. That makes format selection a delivery-policy decision rather than part of segmentation. Keep fallback behavior explicit and test it against the browsers your learners and catalog editors actually use; don't assume that a format choice is universal because it worked in one desktop preview.

The decision table is short enough to use during design review:

Layer Optimize for Reject when
Transparent master edge fidelity and future reuse alpha is empty, dimensions are wrong, or review fixtures regress
Responsive derivative bytes at the required display quality visible halos, lost fine detail, or dimensions exceed the intended slot
Publication manifest atomic, repeatable selection any required rendition is missing or points at another policy version

Lossy derivatives are not suitable when the asset will return to an editing workflow or be composited at unpredictable sizes; keep the master for those uses. Conversely, sending the master to every catalog card wastes bandwidth and couples page performance to an archival choice. Small cards can tolerate a different quality setting than a zoom view. Measure both with representative garments, because dark coats on light backgrounds and translucent fabric expose different edge errors.

Deploy the policy behind a versioned manifest. Generate variants before switching the manifest, verify them, then promote the batch in one state transition. A rollback becomes a pointer change to the previous complete manifest rather than another image-processing run. This also keeps a model or encoder change from producing a half-old, half-new catalog during an upload burst.

Set the threshold, then account for its noise

An oldest-item threshold close to normal processing time detects trouble quickly, but normal variance will page the team during every large batch. A loose threshold protects sleep while allowing real catalog gaps to remain visible longer. There is no universal setting.

Start with the user deadline, subtract the time needed to diagnose and recover, and alert there. Then replay the condition over historical lifecycle events before enabling paging. Examine how many pages would have led to an action: draining a bad policy version, restoring stage capacity, or reconciling a stranded claim. Route non-actionable single-item misses to a ticket and reserve the page for sustained deadline risk across a batch.

Watch the quiet failure too. If no uploads are expected overnight, “no completed jobs” is healthy; if a scheduled catalog import produced accepted records but no transitions, it is not. Pair throughput alerts with a known input signal so absence has context.

False positives have a concrete cost. They train the on-call to wait for a second alert, which erases the early-warning time the threshold was meant to buy. Review every page against the runbook, record whether an action changed the outcome, and tune the window or routing when it did not. Keep the customer deadline fixed while tuning the signal around it.

References

Further reading

Top comments (0)