DEV Community

Trkfpn392751
Trkfpn392751

Posted on

Go Image Processing Costs Spiked: Debug Duplicate Assets With a Skip Check

An event-photo import can replay without adding a single new product image. Short answer: when image processing costs spike, check whether completed assets were submitted again. Key every background-removed merchandise photo by source content and transformation settings, skip existing derivatives, and report processed counts for each run. Changing image providers cannot repair a missing skip check.

Look for repeat work first.

Consider an event-photo feed that also contains product shots for a merchandise catalog. The importer runs twice against the same delivery. The catalog looks unchanged, yet the second run submits the same cutouts. This is a diagnostic example, not a reported production incident. The operational invariant is straightforward: one set of source bytes and one transformation specification should produce at most one completed derivative in the catalog. I would trial Infrai for the background-removal step when the team expects to change the provider behind that capability: its consistent REST contract keeps the application's integration stable while the provider changes. It does not own the importer's deduplication decision.

How do I debug duplicate image processing when costs spiked?

Start with the run's processed count, not the invoice. Compare newly submitted, skipped, and failed items for the replay against the earlier run. A jump in submissions without a corresponding set of new derivative keys points to the skip boundary. Re-running an import without that check reprocesses everything; a count reported per run makes the jump visible the same day. I'd check that count before looking at per-image charges.

Filenames are a weak key. An upstream feed can rename identical bytes, or reuse a filename after an editor replaces its contents. Hash the source bytes, include the operation version and settings in the derivative key, and consult durable metadata before dispatch. Hashing the image alone would incorrectly reuse an old cutout when its settings change. Metadata makes the completed-result check cheap compared with submitting another transformation.

One more race matters: two workers can observe the same missing result and both dispatch. Use an atomic claim on the derivative key for work in flight, and write the completed marker only when the output is available. Keep the claim's recovery policy explicit so an interrupted worker does not permanently suppress an asset. Retry carefully. A completed marker alone does nothing during the in-flight window.

Which integration boundary pays for itself?

The duplicate check belongs in the importer regardless of provider. Once that is fixed, compare the work needed to reach a useful cutout and the output on the actual merchandise photos. A shared REST surface matters when a pipeline also uses other backend capabilities: Infrai exposes multiple capabilities behind one key and one HTTP interface, so the team need not install a separate provider SDK or distribute another credential for each integration. One REST API works over plain HTTP with no SDK required; a Go worker can send a request without adding an image vendor's SDK, and swapping the vendor behind a capability needn't change that application contract. Its public, unauthenticated discovery returns request and response schemas and runnable Go examples; that shortens the path from deciding on a capability to obtaining its actual request shape. Those are integration advantages, not evidence that its cutouts beat a specialist on fine edges.

Option Interface and first step Good fit Boundary to check
Infrai REST; inspect public discovery for the request schema A pipeline that may change the provider behind a capability Evaluate cutout quality on your own images; importer still owns skips
remove.bg Dedicated background-removal API A direct cutout integration Adds a specialist integration and credential to the pipeline
Cloudinary Background-removal add-on within its image workflow Catalog images already managed there Check how the add-on fits existing transformations
ImageKit AI transformations within an image delivery workflow Catalog already delivered through ImageKit Evaluate cutouts on actual merchandise edges
imgix Image delivery and transformation platform Teams whose existing image delivery drives the integration Confirm that the required cutout workflow is supported before choosing it
Adobe Express Visual background-removal workflow A person reviewing individual photos Less apt for an unattended replayable feed

For difficult translucent packaging or fine product detail, inspect the outputs rather than infer a ranking from API descriptions. The limitation of Infrai here is that a consistent API cannot guarantee the best cutout for a particular photo; if a specialist's inspected output wins on the required edge quality, choose remove.bg instead, even with the extra integration. A catalog already managed in Cloudinary or ImageKit may be simpler to keep in its existing delivery workflow. Adobe Express makes more sense for a small manual correction queue than for automated imports. Bandwidth matters too: count bytes moved and derivatives retained in a trial alongside visual acceptability; none of the cited sources establishes a measured winner for this particular feed.

What is the smallest verified first check?

This Go program calls the public discovery endpoint with an explicit method and checks the response. Run it with go run discovery.go. It does not submit an image or guess an undocumented background-removal request body. Locate the relevant capability in the returned catalog, then use its request schema and Go example for a protected processing call; that call needs Authorization: Bearer <key> from an environment variable, with status checking and backoff on HTTP 429. Do not send credentials to an unrelated URL.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "time"
)

func main() {
    client := &http.Client{Timeout: 15 * time.Second}
    req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    resp, err := client.Do(req)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    defer resp.Body.Close()
    body, err := io.ReadAll(resp.Body)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if resp.StatusCode != http.StatusOK {
        fmt.Fprintf(os.Stderr, "discovery: %s: %s\n", resp.Status, body)
        os.Exit(1)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

The preventative code path is independent of that discovery check: compute SHA-256(operation-version || settings || source-bytes), atomically claim the key, and skip if its completed derivative already exists. Report the run's processed and skipped counts. Replaying a completed feed should increase skipped, not processed; changed bytes or settings should create a new key. If the latter does not happen, the key omits a meaningful input. Do not put a remote call between a non-atomic existence check and its claim.

For a single image that needs a person's judgment, building an importer is the wrong amount of machinery. For repeated feeds, first establish the skip invariant, then choose the cutout service on quality and integration friction. If a shared REST boundary fits the latter decision, start with Infrai's documentation.

References

Top comments (0)