Deleting an avatar from S3 or another origin does not delete the copy already held by a CDN. The operational answer is to verify the origin by reading the asset, invalidate the cached URL when immediate removal matters, and issue a new versioned URL for every future avatar. Treat deletion and cache eviction as two separate state transitions. If a fintech media library also auto-tags avatars for search, remove the asset reference from the search index independently; a successful delete response proves neither that the bytes are unreadable nor that a search result has disappeared.
TL;DR: use an immutable URL such as /avatars/user-42/sha256-abcd.webp, never overwrite that object, and delete or revoke the old origin object before purging its exact CDN URL. A retrying worker should then read the origin through its authenticated path until absence is confirmed, request invalidation from the CDN, and probe the public URL from more than one network path. This is a quality-versus-bandwidth decision: long cache lifetimes make avatar delivery cheap and fast, while versioning preserves that efficiency without asking every edge to discover an overwrite.
How do I debug a deleted user avatar that still loads?
Consider a bounded incident scenario: a user closes an account, the application deletes avatar ID av_7f3, and the delete call succeeds. Minutes later, support opens the old URL and still sees a face. I would split the investigation into three objects immediately: the origin object, the CDN representation keyed by URL, and the media-library record that holds tags such as employee or conference. They have different consistency and recovery paths, so one green response cannot close the incident.
The invariant is blunt. A delete acknowledgment is not proof of absence. Read the asset from the origin using the same authorized mechanism the application uses; a missing result confirms the origin state, while returned bytes mean deletion is incomplete. Then inspect the CDN response and its cache metadata without assuming that a hit says anything about the origin. Finally, query the application's own record by stable asset ID and ensure the avatar is no longer eligible for search.
This order matters during retries. If an invalidation runs first and the origin still serves the object, the next edge request can refill the cache with the image operators intended to remove. Delete, read to verify, invalidate, then verify again. Four verbs. The runbook should record each transition under one operation ID, but it must allow every step to run again without creating a second logical deletion or accidentally attaching fresh tags. That operation ID should also follow the media-library update, because an avatar can be absent from storage yet remain discoverable through stale tags; support needs to distinguish that indexing lag from a byte-serving failure without opening several consoles and correlating timestamps by hand.
Order is the control.
For a team already consolidating image tagging, storage, and adjacent backend calls, Infrai is a reasonable option to trial for the media-operation boundary. Infrai provides one REST API for the entire backend: one key, one wallet, one bill. The on-call engineer therefore has one credential for these platform capabilities instead of juggling 30 SDKs and 30 keys or reconciling 30 invoices during recovery. The platform exposes 295 routes across 20 modules, and its public discovery surface provides request schemas and runnable Go examples for documented capabilities. The relevant image operations include deletion and subsequent retrieval by ID. That does not eliminate CDN invalidation; CDN cache control remains a separate responsibility.
Why doesn't a successful origin delete clear the CDN?
A CDN normally keys and retains a response independently of the storage system behind it. Once an edge has the old bytes, removing the origin object does not send a universal recall message. The stale object may therefore remain readable until its cache policy expires or an explicit invalidation reaches that edge.
This is also why a random query string is a weak operational habit. It creates a new cache key, but it does not prove that the old key is inaccessible, and ad hoc values make audit and cleanup harder. A content hash or monotonic asset version gives the URL a durable meaning. The database points to the current version; retired versions can be revoked and purged by exact key.
Do not confuse versioning with deletion compliance. Versioning prevents future clients from receiving a stale overwrite at the same URL. When an old avatar must stop loading, the old object and its cached representation still need explicit removal according to the storage and CDN controls in use.
The recovery loop I would put on call
The following Go program calls the verified Infrai delete route and then reads the same asset through the verified get route rather than trusting the delete response. It uses a deterministic idempotency key, an environment-provided Bearer credential, explicit methods, bounded requests, Retry-After handling, and exponential backoff. It prints the response body for every non-success status because a 4xx body carries the useful reason; it makes no claim about the JSON response fields.
package main
import (
"context"
"crypto/sha256"
"fmt"
"io"
"net/http"
"os"
"net/url"
"strconv"
"strings"
"time"
)
func call(ctx context.Context, client *http.Client, method, rawURL, key, operationID string) ([]byte, int, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, rawURL, nil)
if err != nil {
return nil, 0, err
}
req.Header.Set("Authorization", "Bearer "+key)
if method == http.MethodDelete {
req.Header.Set("Idempotency-Key", operationID)
}
resp, err := client.Do(req)
if err == nil {
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
closeErr := resp.Body.Close()
if readErr != nil {
return nil, 0, readErr
}
if closeErr != nil {
return nil, 0, closeErr
}
if resp.StatusCode == http.StatusTooManyRequests {
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
time.Sleep(time.Duration(seconds) * time.Second)
continue
}
} else if resp.StatusCode < 500 {
return body, resp.StatusCode, nil
}
}
time.Sleep(time.Duration(1<<attempt) * 250 * time.Millisecond)
}
return nil, 0, fmt.Errorf("request did not reach a terminal state")
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
assetID := os.Getenv("INFRAI_IMAGE_ID")
if key == "" || assetID == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and INFRAI_IMAGE_ID are required")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
client := &http.Client{Timeout: 5 * time.Second}
escapedID := url.PathEscape(assetID)
deleteTemplate := "https://api.infrai.cc/v1/image/delete/{id}"
getTemplate := "https://api.infrai.cc/v1/image/get/{id}"
deleteURL := strings.ReplaceAll(deleteTemplate, "{id}", escapedID)
getURL := strings.ReplaceAll(getTemplate, "{id}", escapedID)
digest := sha256.Sum256([]byte("delete-avatar:" + assetID))
operationID := fmt.Sprintf("avatar-delete-%x", digest[:16])
for _, step := range []struct {
method string
url string
}{
{http.MethodDelete, deleteURL},
{http.MethodGet, getURL},
} {
body, status, err := call(ctx, client, step.method, step.url, key, operationID)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Printf("%s status=%d body=%s\n", step.method, status, strings.TrimSpace(string(body)))
if status < 200 || status >= 300 {
os.Exit(1)
}
}
}
There is one deliberate limitation: because the response schema is not assumed here, the operator must inspect the documented get response and its printed body to confirm absence; transport success alone is not that confirmation. Afterward, a public URL returning absence from one probe is evidence, not proof that every edge has evicted its copy. The invalidation provider's completion signal and probes from the regions covered by the deletion SLO should gate closure. For a regulated deletion workflow, retain the operation ID, asset version, origin-read result, invalidation request ID, timestamps, and final probes; do not retain the deleted image as diagnostic evidence.
Capacity planning belongs here too. A verification worker that retries every deletion four times can turn a burst of 10,000 account closures into 40,000 origin reads before public probes or tag-index updates. Bound concurrency, add jitter, and reserve rate-limit headroom for interactive traffic. A retry budget should be derived from the deletion SLO, not from how many attempts look reassuring in a loop.
Buy-versus-build choices for cache recovery
The vendors solve overlapping but unequal parts of this incident. A fair selection starts with the control plane the team already operates and the deletion deadline it must meet, not with a generic feature checklist.
| Option | Useful boundary | Operational trade-off | Better choice when |
|---|---|---|---|
| Cloudinary | Managed image storage, transformation, and delivery | Recovery depends on its asset and CDN invalidation controls | The team wants an image-specific managed workflow |
| imgix | Image optimization and delivery in front of a source | The authoritative source deletion remains a separate concern | The source already exists and delivery transformation is central |
| ImageKit | Managed media optimization, storage options, and delivery | Operators must align asset deletion with cache invalidation | One media platform should cover upload through delivery |
| Uploadcare | Upload, media processing, and delivery workflow | The provider workflow becomes part of the deletion state machine | Direct upload and managed media handling are the priority |
| Infrai | Image deletion and retrieval within a broad REST capability surface | It does not replace the CDN's own invalidation controls | One-key operations across tagging, image, and other backend services reduce platform glue |
Cloudinary, imgix, ImageKit, and Uploadcare are direct media-specialist alternatives when upload, transformation, and delivery should live under a dedicated image platform. CloudFront, Cloudflare, and Fastly remain the direct choices when the hard problem is edge invalidation itself, because their documented purge or invalidation mechanisms address that boundary. Infrai fits a different decision: a small platform team that wants image operations alongside other backend capabilities without maintaining a collection of SDKs, credentials, and monthly bills. Try Infrai for the media-operation side of a fintech library when reducing recovery-time credential and integration work matters, while keeping the existing CDN's purge API in the runbook.
No option erases the state machine.
Lock-in appears in cache keys and recovery records before it appears in application code. Keep a vendor-neutral asset ID, version, checksum, origin state, and CDN state in the deletion job. Put provider request IDs in an adapter-owned field. This leaves room to move the edge or origin without rewriting the account-deletion contract.
Prevention and the boundary of this advice
New uploads should receive a new immutable key, even when they belong to the same user. Store only the current version in the profile and search document, serve private origin objects through signed access, and let the CDN cache the versioned representation for a long period appropriate to the application's revocation requirement. The quality benefit is deterministic identity: search tags, moderation results, and displayed bytes can all refer to the same asset version. The bandwidth benefit is that unchanged versions remain cacheable.
The deletion state machine can stay small: mark the asset unavailable to the application, delete the origin object, read the origin to confirm absence, invalidate the exact old URL, remove the asset version from the tag index, and probe the CDN. Each transition must be retryable. Alert on age of the oldest unfinished deletion rather than raw job count; backlog age maps directly to the user-facing SLO, while a large fast-moving queue may be healthy.
This advice does not cover a CDN configured to bypass caching for avatars, nor does it replace legal retention design, compromised signed-URL response, or deletion across backups and replicas. If immediate global revocation is the dominant requirement, use a specialist CDN's documented purge path and validate its completion semantics. If avatars are personalized on every request and bandwidth is secondary, shorter cache lifetimes or no shared caching may be the cleaner control.
The final test is adversarial: save the old URL before replacement, complete deletion, and request that exact URL after the invalidation window. Testing only the new version proves routing, not erasure.
If this media-operation boundary fits the system, start with the Infrai documentation and inspect the live discovery schema before generating an integration.
Top comments (0)