DEV Community

loganpierce2073
loganpierce2073

Posted on

Property Media Retention Jobs: Verified IDs for Expired Images and Video

Short answer: make the retention worker delete only an object whose expiry record, media type, tenant, and confirmed object ID still agree; keep a small audit record, and treat bandwidth saved by compression as a separate decision from deletion.

For a property-management platform, the bill is usually dominated by bytes served and stored, not by the worker that scans a table. A few thousand listing photos are cheap to inspect; repeatedly serving original 12 MB walkthrough videos to mobile clients is not. The first useful measurement is therefore a bytes-by-media-type report: original image bytes, derived image bytes, video bytes, request count, and cache hit rate. Deleting an expired rendition changes storage, while choosing WebP, AVIF, or an H.264 profile changes delivery bandwidth. Those levers should not be conflated.

Measure first.

For each tenant and media class, take a seven-day baseline of transferred bytes, stored bytes, and cache misses, then project the result against the retention window. A deletion job that removes 10 GB of cold originals may barely move the transfer line, while a 200 MB derivative requested on every listing page can dominate it. Keep the baseline query versioned in source control, because changing a CDN rule halfway through an experiment can make a retention policy look better than it is. I also keep the “bytes avoided” estimate separate from the “bytes deleted” estimate: the former comes from encoding and cache behavior, the latter from lifecycle policy. That distinction makes a finance review much less theatrical and gives operations a number they can reproduce.

What should a retention worker confirm before deleting media?

The worker should receive a candidate record, then re-read the authoritative row inside a transaction. A candidate is eligible only when its expiry timestamp is in the past, its lifecycle state is expired, and the supplied object ID matches the row selected for that tenant. The delete operation must be idempotent: a retry after a timeout should produce the same final state as the first successful attempt.

This is a practical exactly-once mindset built on at-least-once delivery. A queue can deliver a job twice. The database can record deleted_at before an object-store response arrives, or after it arrives. Use an operation key such as (tenant_id, object_id, retention_version) and write an append-only audit event for every decision: eligible, skipped_mismatch, deleted, or delete_failed. The audit trail is more valuable than a pretty dashboard when a landlord asks why a lease video disappeared.

Here is the core decision in Go. The storage interface deliberately returns an “already absent” result as success, because absence is the desired terminal state.

package retention

import (
    "context"
    "time"
)

type Media struct {
    TenantID   string
    ObjectID   string
    ExpiresAt  time.Time
    State      string
    Generation int64
}

type Store interface {
    Find(ctx context.Context, tenantID, objectID string) (Media, error)
    Delete(ctx context.Context, tenantID, objectID string) (alreadyAbsent bool, err error)
    Audit(ctx context.Context, tenantID, objectID, outcome string) error
}

func Expire(ctx context.Context, s Store, tenantID, objectID string, now time.Time) error {
    m, err := s.Find(ctx, tenantID, objectID)
    if err != nil {
        return err
    }
    if m.TenantID != tenantID || m.ObjectID != objectID || m.State != "expired" || m.ExpiresAt.After(now) {
        return s.Audit(ctx, tenantID, objectID, "skipped_mismatch")
    }

    _, err = s.Delete(ctx, tenantID, objectID)
    if err != nil {
        _ = s.Audit(ctx, tenantID, objectID, "delete_failed")
        return err
    }
    return s.Audit(ctx, tenantID, objectID, "deleted")
}
Enter fullscreen mode Exit fullscreen mode

The example omits the transaction adapter because its boundary depends on the database, but the invariant is not optional: the row lookup and the state transition must be protected from a concurrent renewal or replacement upload. Never trust an object ID copied from a client request; the worker should consume an internal, signed job payload and still verify it against current data.

How do image and video policies change the bandwidth trade-off?

Images usually have several derivatives: a thumbnail for search, a medium card image, and an original for inspection. A property manager may retain the original for a legal or maintenance period while expiring the medium derivative sooner. Video has a different cost curve: a poster frame can remain after the source video expires, but a streaming manifest that references missing segments creates a broken listing. Retention must understand the relationship, not delete each key independently.

Compression belongs before this worker. Use content negotiation and a measured quality ladder; compare perceptual quality against bytes at the actual viewport sizes. MDN’s media-format guidance is a useful standards-oriented starting point, but a format that decodes on a developer laptop may still be a poor choice for an older inspection tablet. Your mileage may vary.

I once assumed that deleting the largest objects would produce the biggest saving. The report was wrong because cache misses, not object size, drove the monthly transfer spike. The corrective job kept frequently viewed derivatives and expired only orphaned originals after confirmation. That change reduced risk, while the bandwidth decision came from resizing and encoding tests.

What failure modes belong in the worker’s design?

The dangerous cases are ordinary: a lease is renewed while a stale job is running, an upload reuses an object key, or a timeout hides a successful delete. Version the retention record, include the generation in the job, and make a mismatch a recorded no-op. Do not infer success from a missing object alone; record the check and preserve the tenant boundary.

Keep deletion and serving decoupled. A serving path should tolerate a missing derivative by selecting an approved fallback, while never resurrecting an expired original. Metrics should include eligible candidates, mismatch skips, idempotent repeats, delete latency, and bytes removed. Alerts belong on a rising mismatch rate or audit-write failure, not on a single absent object.

Compliance changes the retention window. Privacy rules, contractual lease terms, litigation holds, and regional storage requirements can override a default TTL. The worker needs a hold check and an operator-visible reason for every extension. The catch is that a short retention window is not automatically safer if it destroys evidence required by a local rule; get that policy reviewed before tuning a queue.

Where does this approach stop being a good fit?

An internal worker is a poor fit when media is shared across many independent tenants without a reliable ownership key, when legal holds are managed outside the system, or when the storage provider cannot offer an auditable delete result. In those cases, keep the source under a governed archive and use a dedicated lifecycle service with reviewable policies. Stick with a simpler scheduled sweep only for disposable, single-tenant derivatives where a mismatch cannot affect a customer record.

The design also does not solve visual quality by itself. If the product requirement is “the listing must look sharp on every device,” run representative image and video tests, publish a quality budget, and measure transfer bytes alongside engagement. Retention is the final gate on an object’s life, not a substitute for an encoding strategy.

References

Top comments (0)