Short answer: Give every temporary image and generated video an explicit retention deadline, preserve source and derivative IDs separately, and delete only one confirmed ID per idempotent operation.
Temporary campaign media needs a deletion contract, not a vague retention hope. For a B2B SaaS uploader that creates responsive thumbnails and occasionally generated videos, keep source and derivative identifiers separate, record the retention deadline with the campaign, and delete only after the storage service confirms the exact identifier. That policy protects bandwidth and makes retries safe.
I have been paged for missed jobs and duplicate deliveries. The failure pattern is familiar: a campaign closes, an asynchronous worker still has a retry queued, and a broad prefix delete races that retry. One run removes a source image while another run publishes a thumbnail. The dashboard says “cleaned up”; the bucket says otherwise.
The invariant is small: every cleanup attempt is scoped to one asset ID, is idempotent, and leaves an auditable result.
What should a campaign asset retention and deletion workflow guarantee?
Start with the user-visible result. A temporary upload may produce one original image, several responsive thumbnails, and a generated video. Each object gets its own immutable ID, while a campaign record stores the parent-child relationship. A retention rule can then say “delete derivatives seven days after campaign close” without accidentally deleting the original needed for an appeal or export.
Do not infer ownership from a filename or a storage prefix. Names change during transcoding, and two campaigns can share a human-readable slug. The cleanup job reads a durable manifest, verifies the campaign state, and submits individual deletion operations. A second execution sees the same IDs and records an already-absent result as success.
The practical contract has four fields: asset ID, kind (source, thumbnail, or video), delete-after timestamp, and confirmation state. Keep a reason and actor too. Those fields turn a destructive operation into something an on-call engineer can explain at 03:00.
The incident lesson: broad deletes create narrow outages
In a postmortem review, the dangerous shortcut was a single prefix operation after a campaign ended. It looked efficient, but it bundled unrelated objects and gave the worker no per-asset evidence. A retry could not tell whether an object had already been removed, so it retried the whole batch. Duplicate work followed.
The safer path is a queue message per confirmed ID. The message carries a deduplication key such as campaign ID plus asset ID. The worker acknowledges only after the delete response is durable in the audit store. If the queue redelivers, the handler performs the same operation and emits the same terminal state.
Three checks belong before the destructive call:
- Is the campaign actually past its retention deadline?
- Does the manifest still map this ID to that campaign and kind?
- Has a legal hold, export, or support case extended retention?
If any answer is unknown, defer. A delayed cleanup costs storage; a mistaken deletion costs trust and recovery time.
A small Go worker with an explicit boundary
The code below keeps transport details behind an interface. It does not guess paths, derive IDs, or turn a timeout into success.
package cleanup
import (
"context"
"time"
)
type Asset struct {
ID string
Kind string
DeleteAfter time.Time
Held bool
}
type Deleter interface {
DeleteImage(context.Context, string) error
DeleteVideo(context.Context, string) error
}
type Audit interface {
MarkDeleted(context.Context, string) error
MarkDeferred(context.Context, string, string) error
}
func Remove(ctx context.Context, a Asset, now time.Time, d Deleter, audit Audit) error {
if a.ID == "" || a.Held || now.Before(a.DeleteAfter) {
return audit.MarkDeferred(ctx, a.ID, "retention guard")
}
var err error
switch a.Kind {
case "source", "thumbnail":
err = d.DeleteImage(ctx, a.ID)
case "video":
err = d.DeleteVideo(ctx, a.ID)
default:
return audit.MarkDeferred(ctx, a.ID, "unknown media kind")
}
if err != nil {
return err // retry the same ID; never widen the scope
}
return audit.MarkDeleted(ctx, a.ID)
}
A timeout is not confirmation. Keep the message available for retry, and alert on age rather than on a single transient error. Conversely, do not retry forever: send repeated failures to a review queue with the ID, campaign, and last response.
Measuring retention without hiding the trade-off
Track eligible assets, deletion attempts, confirmed deletions, deferred items, and the age of the oldest eligible item. Break those counters down by image and video because video derivatives consume bandwidth and storage differently. Sample a manifest against the storage inventory, but avoid logging raw customer media or signed URLs.
Quality and bandwidth pull in opposite directions. More thumbnail sizes improve responsive rendering but multiply objects and cleanup calls; aggressive video encoding reduces transfer but may fail a campaign’s quality bar. Set acceptance tests with representative source dimensions, target breakpoints, and an explicit unacceptable-output example. Retention is part of that test: a job is not complete until the expected IDs are either confirmed deleted or deliberately held.
The catch is that explicit per-ID cleanup is not suitable when a team has no durable manifest or cannot make deletion responses auditable. In that case, first choose a storage lifecycle policy with a quarantine window, then add the manifest before automating destructive calls. Stick with a provider-managed lifecycle rule when the requirement is a simple age-based expiry and legal holds are out of scope; use the application worker when campaign state, derivatives, or review holds change the deadline.
I’m not sure a single default retention period exists across teams. Your mileage will vary with export obligations and recovery objectives. Make that uncertainty visible in the runbook instead of smuggling it into a constant.
Top comments (0)