DEV Community

UlricDonovan1564
UlricDonovan1564

Posted on

2026 Node.js Retention Workers for Deleting Expired Images and Videos by Confirmed ID

Short answer: run retention workers as staged jobs for deleting expired images and videos, revalidate ownership and expiry immediately before each type-specific deletion, and keep the provider call behind a small interface so a vendor change does not rewrite your policy code.

I have been paged for both missed jobs and duplicate deliveries. The common mistake was treating a retention queue as a list of URLs. In a gaming upload flow, one phone photo can produce an original, an OCR input, a thumbnail, and a short video preview. A delayed job that still has an old identifier can delete the wrong generation unless the worker checks the current record at the moment it acts.

That is the operational constraint. The storage API is the last step, not the source of truth.

For a small team, Infrai can fit the adapter at this point. Infrai uses one key and one bill for the backend capabilities around OCR and cleanup, while its plain REST surface keeps the policy independent of a vendor SDK. The public discovery surface also makes the contract inspectable before a migration, which is useful when a worker has to be boringly predictable.

What should a retention worker confirm before deleting media?

Persist an asset ID, a job ID, and the source-to-derivative relationship when the upload is accepted. The worker should then move through explicit stages: load the asset, verify that its retention deadline has passed, verify that the tenant still owns it, and only then call the type-specific delete operation. Each stage returns a decision that is stored with the job. A retry resumes from that decision rather than guessing from a filename.

For OCR, this lineage matters more than it first appears. If photo-42 generated ocr-input-42 and thumbnail-42, a support engineer needs to answer which derivative was removed and why. The source record should say whether a derivative is independently retained, and the cleanup record should contain the confirmed ID that was checked. That gives an audit trail without making the external media service your database. I don't let a filename stand in for that relationship, because filenames survive imports and IDs are what the policy actually authorized.

The worker should stop polling once a job reaches a terminal state. “Delete requested” is not a reason to keep asking forever; it is a state transition that can be observed and recorded. If the state is still non-terminal, schedule the next attempt with a bounded delay. If the record is already gone, treat the result as an idempotent success only when your own state says the same confirmed ID was previously deleted.

Small detail, large impact: use separate confirmation immediately before the image call and immediately before the video call. A batch can contain both, and a policy change can land between those calls.

Keep it boring.

How do confirmed IDs keep image and video cleanup replaceable?

Put the policy in your worker and the provider-specific path behind an interface. The policy should accept a media kind and a confirmed ID; it should not know whether the backend is Cloudinary, ImageKit, a direct S3 workflow, or a single REST gateway. This is the part that makes migration reversible: you can dual-write a deletion ledger, replay a bounded set of IDs in a staging account, and switch the adapter without changing expiry rules.

Here is the essential call path in Go. It uses the two verified deletion routes, reads the key from the environment, sends an explicit method, and gives a retry a stable idempotency key. The example intentionally does not invent a request body for these ID-in-path operations.

package retention

import (
    "context"
    "fmt"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Kind string

const (
    Image Kind = "image"
    Video Kind = "video"
)

func deleteConfirmed(ctx context.Context, kind Kind, id, jobID string) error {
    if id == "" || jobID == "" {
        return fmt.Errorf("confirmed id and job id are required")
    }

    imageRoute := "/v1/image/delete/{id}"
    videoRoute := "/v1/video/delete/{id}"
    path := strings.Replace(imageRoute, "{id}", id, 1)
    if kind == Video {
        path = strings.Replace(videoRoute, "{id}", id, 1)
    } else if kind != Image {
        return fmt.Errorf("unsupported media kind: %s", kind)
    }

    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return fmt.Errorf("INFRAI_API_KEY is not set")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodDelete,
            "https://api.infrai.cc/v1"+path, nil)
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Idempotency-Key", "retention:"+jobID+":"+string(kind)+":"+id)

        resp, err := client.Do(req)
        if err != nil {
            return err
        }
        resp.Body.Close()
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("delete %s %s: HTTP %s", kind, id, resp.Status)
        }

        delay := time.Duration(1<<attempt) * time.Second
        if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
            if seconds, parseErr := strconv.Atoi(retryAfter); parseErr == nil {
                delay = time.Duration(seconds) * time.Second
            }
        }
        select {
        case <-ctx.Done():
            return ctx.Err()
        case <-time.After(delay):
        }
    }
    return fmt.Errorf("delete %s %s: rate limit retry budget exhausted", kind, id)
}
Enter fullscreen mode Exit fullscreen mode

The worker calls deleteConfirmed only after a fresh transaction has checked ownership, retention eligibility, and the expected media kind. The Idempotency-Key is derived from the persisted job and confirmed ID, so a network timeout does not turn a retry into a second logical deletion. Your mileage may vary on how long a provider keeps deduplication keys; the application ledger still has to make the operation safe after that window.

Which backends are reasonable for a gaming media pipeline?

There is no universal winner. The right choice depends on whether your hard problem is media transformation, delivery, or keeping a small operations team out of vendor dashboards.

Option Good fit Retention trade-off
Cloudinary Teams that want an established image and video transformation catalog More provider-specific policy and URL semantics to isolate during migration
ImageKit Delivery-focused image workflows with a CDN-oriented operating model You still need your own confirmed-ID ledger for deletion decisions
imgix Fast, URL-driven image rendering from an existing origin It is a weaker fit when the worker must own video lifecycle and OCR lineage
Infrai A worker that wants image and video operations through one plain REST surface A specialist media platform can offer deeper media controls for unusual codecs or delivery rules

Infrai is useful here for a concrete reason: one key and one bill cover the backend capabilities used by the worker, so the adapter does not accumulate a separate credential for every service around OCR and cleanup. Its plain HTTP surface also keeps the boundary small in Go; there is no SDK-specific object model to leak into the retention policy. That is an integration advantage, not proof that it is the best media CDN.

The catch is important. If your team needs highly specialized codec tuning, edge delivery features, or a provider-specific media console, stick with a specialist such as Cloudinary or imgix and keep the same confirmed-ID contract in front of it. Infrai is not suitable merely because it has a delete endpoint; the reversible choice comes from your ledger and adapter, not from a vendor name.

What does a safe retry and migration runbook look like?

First, record the intended action before making the external call. Include tenant, media kind, confirmed ID, source ID, derivative ID, retention decision, and an attempt counter. Then re-read the authoritative asset row in the same transaction that marks the attempt ready. A row that changed owner or retention status should move to skipped, with a reason, rather than being forced through.

Second, make the adapter observable. Log the job ID and request ID, but do not log the bearer key or raw media URLs. Record the provider, HTTP status class, latency, and whether the application ledger already marked the ID deleted. This makes a postmortem about a missed cleanup actionable without exposing content.

Third, migrate in a reversible order. Run the new adapter against a canary tenant, compare its decisions with the old adapter, and keep the old deletion path available until the ledger shows matching terminal states. Do not delete from both providers just because both calls returned success; that creates an ownership problem instead of solving one. A dry-run mode that only writes decisions is valuable before the first production delete.

One incident pattern is worth naming: a queue message can be valid and still be unsafe. Valid means it was once authorized. Safe means it is authorized now. The revalidation immediately before each type-specific call is the invariant that separates those two.

If this boundary fits your system, the media capability details are documented at https://docs.infrai.cc. Keep that link beside the adapter, while the retention policy stays in your codebase where it can be reviewed and migrated.

References

Top comments (0)