A deletion request is complete only when you can prove that every playable or viewable copy is gone, or that a documented retention rule still applies. For an edtech product, I would process the request asynchronously, freeze new derivatives, and issue an evidence record that follows the asset through storage, thumbnails, moderation queues, exports, and caches.
Short answer: delete on demand, but verify against an asset lineage graph rather than trusting one successful object-storage response. Upload-time moderation and deletion are separate decisions; combining them creates gaps.
The decision note: what must be true before you call an asset erased?
| Decision | Prefer | Why |
|---|---|---|
| When to moderate an uploaded image | At upload, before publication | Prevents an unsafe image from becoming visible or copied into derivatives |
| When to execute an erasure request | On demand, through a durable job | A request can arrive long after upload and must cover every copy |
| What counts as proof | Lineage plus independent verification | A 204 response says a request was accepted, not that a thumbnail or cache disappeared |
The revenue-per-hour test matters here. A synchronous delete endpoint looks cheap to build, but it makes a partial outage indistinguishable from success and encourages a support ticket when a job stalls. A queue, an idempotency key, and a small audit table take longer on day one. They also let a one-person team ship weekly without manually inspecting buckets.
I would keep moderation in the upload path because the user is waiting to publish. I would keep erasure out of that path because a deletion request is a workflow with retries, deadlines, and evidence. Two clocks.
Ship it.
How should an edtech app verify media erasure requests across image and video assets?
Start with identity, not filenames. Assign an immutable asset ID when the original arrives and record a parent ID for every derivative: resized image, extracted video frame, preview GIF, subtitle file, and moderation copy. The graph can live in a relational table; it does not need a specialized database. What matters is that a query can answer “show me every descendant of asset A” at a known point in time.
The delete worker then walks that graph. It marks the request as running, issues idempotent deletes to each storage class, invalidates application references, and records the provider response. A second pass asks the storage layer for metadata or a bounded read and compares the result with the expected absence. For a CDN, verification means checking that the URL is no longer served from the application path and that the purge operation is recorded. Do not call a cache purge proof if an origin URL still resolves.
Here is the shape of the worker. The interfaces are deliberately boring so they can wrap local disk, object storage, or a test double.
type Asset = {
id: string;
parentId: string | null;
key: string;
kind: "image" | "video" | "derivative";
};
type ErasureEvidence = {
assetId: string;
checkedAt: string;
observed: "absent" | "present";
operationId: string;
};
interface MediaStore {
remove(key: string, operationId: string): Promise<void>;
exists(key: string): Promise<boolean>;
}
export async function eraseTree(
root: Asset,
descendants: Asset[],
store: MediaStore,
operationId: string,
): Promise<ErasureEvidence[]> {
const evidence: ErasureEvidence[] = [];
const targets = [root, ...descendants];
for (const asset of targets) {
await store.remove(asset.key, operationId);
const stillThere = await store.exists(asset.key);
evidence.push({
assetId: asset.id,
checkedAt: new Date().toISOString(),
observed: stillThere ? "present" : "absent",
operationId,
});
}
return evidence;
}
A present observation is not a hidden retry. It is a failed request state that needs an owner and an alert. The API should return a request ID, not pretend that the whole tree vanished during one HTTP call.
The failure modes that make a green delete response misleading
The first trap is derivative drift. A moderation service may create a 320px thumbnail while a video pipeline writes a poster frame under a different prefix. If the derivative table is populated after the copy is created, a request racing that write can miss it. Freeze derivative creation as soon as the request is accepted, then reconcile newly discovered children before closing the request.
The second trap is retention ambiguity. Backups, legal holds, and replicated storage can have different lifetimes. “Deleted from the primary bucket” is a narrow statement. Your policy should say which systems are in scope, how long immutable backups remain, and how access is blocked during that period. The catch is that a strict immediate-erasure promise is not suitable when your backup platform cannot selectively remove an object; in that case, document the retention exception and encrypt backups with a key you can revoke.
The third trap is identifier reuse. Never recycle an asset ID or object key after deletion. A later student upload that happens to use the same filename must not inherit an old audit record.
I once treated a 404 from a thumbnail URL as proof. It was only proof that one route stopped serving bytes. The original still appeared in an export job because that job read from a separate manifest. That mistake cost a release window, not because deletion is exotic, but because the evidence model was too small.
A practical evidence record for a one-person team
Keep one append-only record per operation with the requester, policy version, asset-root ID, discovered descendants, attempted locations, verification observations, and timestamps. Hash the manifest before writing it to durable storage. The hash does not make data disappear; it makes later tampering easier to detect.
Use metrics that describe work, not vanity throughput: age of the oldest pending request, percentage verified within the policy deadline, descendants discovered after the first pass, and count of requests blocked by a legal hold. Sample the full chain in staging with real format variety. MDN’s media format guidance is a useful reminder that containers and codecs change how you generate previews and test playback, even when the deletion contract stays the same.
Keep logs free of signed URLs and raw image bytes. A request ID and asset IDs are enough to join the audit trail. Your support dashboard should show “awaiting verification” as a real state, with a retry button that reuses the same idempotency key. Boring controls win.
When on-demand erasure is the wrong fit
On-demand jobs are a poor fit for a live classroom where a teacher expects an image to disappear before the next student refreshes. Add an immediate visibility flag in the application database, then let the durable erasure job finish storage cleanup. Conversely, upload-time deletion is wasteful for archives with a strict legal hold or for assets that must be retained for grading; apply the hold before enqueueing the destructive work.
Stick with a simpler synchronous path when the asset has exactly one storage location, no derivatives, and a measured request volume small enough for a human review queue. Once images and videos fan out to previews, exports, and caches, the lineage graph and evidence record pay for themselves. Your mileage may vary because provider verification APIs differ; test the exact semantics you rely on before promising a deadline.
The goal is not a green checkmark. It is a defensible answer to “where did this media go?” that your future self can produce in minutes, even during a busy launch week.
Keep the first version small: one graph query, one queue, one verifier, and one dashboard. Add complexity only when a new storage location appears in the lineage.
Top comments (0)