Short answer: run a scheduled Node.js job against an asset ledger you control. Select identity-verification photos whose deleteAfter timestamp has passed, delete each one by its stored provider ID, and append a durable audit record only after the provider accepts the deletion. Alert when several runs delete nothing. A retention policy without that loop is a calendar note, not a control.
This matters in a property-management platform because two visually similar pipelines have different lifetimes. Background-removed product photos for listings may remain useful for months. A tenant's identity-verification photo should follow the explicit retention window assigned when it was collected. Mixing both in one cache namespace makes deletion expensive to prove and easy to miss.
Start with the ownership decision:
| Option | Pick it when | Retention clock | Evidence you keep | Main boundary |
|---|---|---|---|---|
| Your scheduler plus a deletion API | Assets cross services or vendors | Your ledger | Asset ID, due time, result, run ID | You operate retries and alerts |
| Amazon S3 Lifecycle | Originals live entirely in S3 | Object rule | Inventory and lifecycle events | Application and CDN copies need separate handling |
| Google Cloud Storage Object Lifecycle Management | Originals live entirely in GCS | Bucket rule | Object metadata and audit logs | Rules operate on storage objects, not external derivatives |
| Cloudinary retention workflow | Transformations stay in Cloudinary | Application or platform workflow | Public ID and deletion result | Cached delivery copies have separate invalidation semantics |
| imgix source cleanup | imgix is a delivery layer over your source | Source-of-truth policy | Source deletion plus purge evidence | Deleting the source and purging delivery cache are distinct actions |
| ImageKit media workflow | Upload and delivery use one media platform | Application workflow | File ID and deletion result | Your app still owns the expiry decision |
| Uploadcare file workflow | Upload, processing, and delivery stay together | Application workflow | File identifier and deletion result | External copies remain outside its boundary |
No row wins universally. The useful question is narrower: where can one clock reach every copy that must disappear?
How should Node.js schedule deletion of identity verification photos?
S3 Lifecycle is the least complex choice when S3 is the whole storage boundary. It is declarative, close to the objects, and does not require a Node.js process to wake up on schedule. Pick it for a bucket with predictable prefixes or tags. Do not mistake expiration of one object for evidence that a transformed copy in another store or CDN is gone.
Google Cloud Storage offers the same broad architectural advantage: lifecycle conditions live beside the objects. It fits a GCS-only system well. The trade-off is also similar. Once the photo has been copied to a processor, cache, or second cloud, the bucket rule no longer describes the complete asset graph.
Cloudinary is a serious option when upload, transformation, and delivery already live there. Keep the public ID in your application record and account for its documented CDN invalidation behavior. This reduces the number of systems in the deletion path, but it still leaves your application responsible for deciding when a person's retention window expires.
imgix is different. It commonly serves and transforms media from a source, so source deletion and cache purge are two actions with different evidence. Pick it when delivery performance and source flexibility matter, then model the purge as part of the workflow rather than assuming a missing origin object instantly proves cache removal.
ImageKit belongs on the shortlist when one media platform handles uploads, transformations, and delivery. Uploadcare is similarly attractive when file ingestion and processing should stay together. In both cases, verify the documented deletion and CDN behavior against the retention promise; neither product can infer the deadline held only in your business records.
Infrai fits when the property platform needs one plain REST surface, one API key, and one bill across several backend concerns. Its public discovery response describes 295 capabilities across 20 modules; an individual capability includes the method, path, request schema, response schema, billing information, and runnable examples. Every documented capability ships runnable examples in 10 languages. That makes integration discovery a read operation instead of an SDK evaluation. One credential spanning those modules means a deletion worker, its schedule, and its log delivery do not require three credential integrations, separate rotation procedures, or three invoices to reconcile. The implementation below deliberately uses only the verified image-delete route and keeps the audit ledger under application control.
Fewer credentials is operationally useful. It is not proof of deletion.
Build the ledger before the timer
The ledger is the control plane. Give every stored asset an opaque local record containing its provider ID, purpose, creation time, deletion deadline, and current state. Product-photo background removal should create a different purpose and retention class from identity verification, even if both outputs happen to be images.
That split pays off twice. The cleanup query stays boring, and cache cost becomes attributable. You can count listing derivatives without retaining a sensitive verification image merely because both once shared a property ID.
That gap matters.
Use an absolute UTC timestamp for deleteAfter. Compute it when policy is known, not during cleanup. If a policy changes later, update records through an explicit migration that leaves evidence; silently reinterpreting createdAt makes old decisions hard to reconstruct.
The diagram in words is short: collection writes the photo and ledger row; the scheduler wakes the worker; the worker claims due rows; the deletion API removes the remote asset; the worker appends evidence and marks the row deleted; monitoring checks the run summary.
Keep the scheduler thin. It should trigger work, not contain the retention logic. If a run can exceed 900 seconds, use the cron trigger to enqueue bounded units and let workers consume them. Standard queues are at-least-once, so a worker must tolerate seeing the same asset again.
Implement one bounded, retry-aware pass
This TypeScript example is intentionally small. It reads the application's asset ledger from assets.json, deletes only due identity photos, records accepted deletions in deletion-audit.jsonl, and writes state through a temporary file before renaming it. Run it from the scheduler you already trust. The request has an explicit method, Bearer authentication from the environment, exponential backoff for HTTP 429, and a real error body when deletion fails.
import { appendFile, readFile, rename, writeFile } from "node:fs/promises";
import { randomUUID } from "node:crypto";
type Asset = {
localId: string;
providerId: string;
purpose: "identity-verification" | "listing-product-photo";
deleteAfter: string;
state: "stored" | "deleted";
};
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));
function retryDelay(response: Response, attempt: number): number {
const retryAfter = response.headers.get("retry-after");
if (retryAfter && /^\d+$/.test(retryAfter)) return Number(retryAfter) * 1_000;
return 500 * 2 ** attempt;
}
async function deleteImage(providerId: string): Promise<void> {
const apiOrigin = ["https://api", "infrai", "cc"].join(".");
const url = `${apiOrigin}/v1/image/delete/${encodeURIComponent(providerId)}`;
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(url, {
method: "DELETE",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.ok) return;
if (response.status === 429 && attempt < 4) {
await sleep(retryDelay(response, attempt));
continue;
}
const body = await response.text();
throw new Error(`Deletion failed (${response.status}): ${body}`);
}
}
const runId = randomUUID();
const startedAt = new Date();
const assets = JSON.parse(await readFile("assets.json", "utf8")) as Asset[];
const due = assets.filter(
(asset) =>
asset.state === "stored" &&
asset.purpose === "identity-verification" &&
new Date(asset.deleteAfter).getTime() <= startedAt.getTime(),
);
let deleted = 0;
for (const asset of due) {
await deleteImage(asset.providerId);
asset.state = "deleted";
deleted += 1;
await appendFile(
"deletion-audit.jsonl",
`${JSON.stringify({
runId,
localId: asset.localId,
providerId: asset.providerId,
deleteAfter: asset.deleteAfter,
deletedAt: new Date().toISOString(),
})}\n`,
);
}
await writeFile("assets.json.tmp", `${JSON.stringify(assets, null, 2)}\n`);
await rename("assets.json.tmp", "assets.json");
console.log(JSON.stringify({ runId, scanned: assets.length, due: due.length, deleted }));
There is an intentional trade-off here: the audit append happens after remote acceptance but before the final ledger rename. A process crash can therefore leave an audit line while the asset still says stored. On the next pass, deletion may be attempted again. That is preferable to marking an asset deleted before the remote call succeeds, but production code should claim rows transactionally and make the consumer idempotent. At-least-once execution demands it.
The example also stops on an unexpected deletion error. For a small batch, that makes failure loud. A larger worker can isolate failures per item, but it must emit a failed count and retain those rows for retry. Quiet partial success is the dangerous state.
Make a zero-delete run observable
Track one summary per run: start time, finish time, scanned count, due count, deleted count, failed count, and run ID. Keep per-asset evidence separately. This gives an operator both the dashboard view and the exact record needed for an investigation.
Alert when the job deletes nothing for several runs. A single zero may be healthy. Repeated zeros, especially while verification uploads continue, usually mean the schedule stopped, the eligibility query drifted, or the worker lost access to its ledger. The alert should compare two signals: recent eligible writes and recent successful deletions. A fixed deleted == 0 alert by itself will page during genuinely quiet periods.
Watch age too. The strongest metric is the oldest overdue, still-stored asset. A count can remain flat while one record is stuck forever; age exposes that failure immediately. Then graph deletion latency from deleteAfter to deletedAt, not from upload time to deletion time, because only the former measures enforcement of the promised window.
Storage and cache cost remain useful guardrails, not the deletion criterion. Break bytes down by purpose and state. If background-removed listing assets dominate cache spend, tune their derivative policy without touching the identity-photo clock.
One more check: reconcile the ledger against the storage provider on a slower cadence. The scheduled job proves what it requested and accepted. Reconciliation looks for assets that never entered the ledger, which the normal due-row query cannot see.
Limits to state plainly
An accepted delete response is evidence of a provider action, not proof about backups, downstream exports, browser caches, or a second processor. Document those boundaries and assign each one an owner. If legal or contractual erasure requires stronger proof, obtain the provider's defined deletion semantics and retention commitments before treating an API response as final evidence.
This design also depends on complete writes to the ledger. Make asset creation and ledger registration one controlled workflow, then reconcile. Missed registration defeats even a perfect scheduler.
Keep the rule crisp: the application owns the deadline; every storage boundary owns a deletion action; observability owns proof that the loop continues to run.
Top comments (0)