Short answer: schedule deletion from a durable record of each identity verification photo's expiry time, then make the worker idempotent and observable. Do not derive eligibility from the upload timestamp at run time. The upload path should write the object and its retention record together; a separate Node.js worker can delete expired objects in small, retryable batches.
This is a B2B SaaS workflow, so the hard part is not calling delete. The hard part is proving that a photo was eligible, that the right object was targeted, and that a retry cannot delete a newer upload with a recycled key.
How should Node.js schedule deletion of identity verification photos after a retention window?
Treat retention as data, not as a cron expression. For every uploaded image, persist an immutable object key, a tenant identifier, a creation time, an expiry time, and a deletion state. expiresAt is calculated once from the policy in force when the image is accepted. If a tenant changes its policy tomorrow, yesterday's evidence does not silently move its deadline.
The object key needs a generation component. A key such as tenant-42/idv/front.jpg is dangerous: a late retry could remove a replacement image. A generated key such as tenant-42/idv/9b3.../front.jpg gives the deletion job an exact target. Keep the database row after deletion, with a tombstone and timestamps. It is useful evidence when a customer asks what happened.
Keys are part of the safety boundary.
The processing choice is simple. Delete at upload time when the retention window is zero or when a failed moderation decision must never leave a file behind. Delete on demand with a scheduled worker for normal windows, because the worker can pace requests, retry transient storage errors, and report lag. These are different guarantees; calling both paths “automatic deletion” hides an important operational distinction.
Here is the smallest working shape. The storage interface is deliberately boring: an adapter can map it to object storage, a filesystem in a test, or an internal media service without changing the retention logic.
type PhotoRecord = {
id: string;
tenantId: string;
objectKey: string;
expiresAt: Date;
state: "active" | "deleting" | "deleted" | "failed";
deleteAttempts: number;
};
interface PhotoStore {
claimExpired(limit: number, now: Date): Promise<PhotoRecord[]>;
deleteObject(objectKey: string): Promise<void>;
markDeleted(id: string, deletedAt: Date): Promise<void>;
markFailed(id: string, reason: string): Promise<void>;
}
export async function deleteExpiredPhotos(
store: PhotoStore,
now = new Date(),
limit = 100,
): Promise<number> {
const rows = await store.claimExpired(limit, now);
let completed = 0;
for (const row of rows) {
try {
await store.deleteObject(row.objectKey);
await store.markDeleted(row.id, now);
completed += 1;
} catch (error) {
const reason = error instanceof Error ? error.message : "unknown error";
await store.markFailed(row.id, reason);
}
}
return completed;
}
claimExpired must be an atomic state transition, usually from active or failed to deleting, guarded by expiresAt <= now. Two workers may start at the same time, but only one should claim a row. The delete operation itself should be idempotent: an already absent object is a successful end state, not a reason to resurrect the record. If the storage adapter distinguishes “absent” from other errors, normalize only the absent case.
The failure modes that decide the design
Clock drift is the first trap. Use UTC timestamps and compare instants, never local date strings. A daylight-saving transition should not change a 30-day retention window. Inject now into the worker so tests can pin the boundary at one second before and one second after expiry.
The second trap is partial success. A database transaction cannot usually include a remote object delete. That is fine, as long as the states describe the gap. A row in deleting is recoverable after a worker restart; a periodic reaper can reclaim it after a lease timeout. A row in deleted means the object is no longer expected to exist, even if a later audit finds an unrelated object with a different generation key.
The third trap is a queue that quietly falls behind. Measure the oldest expiresAt among active rows, the count of rows in deleting, and the age of the last successful batch. A daily job with a 48-hour lag is not meeting a 24-hour deletion promise. Alert on lag, not just on process exit status.
That metric caught a subtle incident in one of my test harnesses: the process exited cleanly after claiming rows, but a worker crashed between the remote delete and the tombstone write. The next run saw deleting rows and skipped them forever because the query only selected active. The fix was not a bigger retry count. I added a lease timestamp, made expired leases claimable, and wrote a test that kills the worker at each boundary. Now a restart can reclaim the row, an absent object counts as success, and the audit event records the second attempt without pretending the first attempt never happened. The extra columns look fussy until somebody asks for a deletion report on a Friday afternoon.
Exactly.
The test is intentionally unpleasant. It freezes the clock at 12:00:00Z, claims a row with a ten-minute lease, advances time to 12:10:01Z, and starts a second worker while the first worker is still paused. The second worker may reclaim the row, but it must still target the same immutable object key. Next, the first worker resumes and receives an idempotent “already handled” result from the storage adapter. Finally, the harness checks that there is one tombstone, two attempt records, and no active row. This sequence is longer than the happy path, yet it is the path that tells me whether a deploy, a process kill, or a delayed network response can turn a routine privacy job into an ambiguous one. I keep this test next to the adapter contract, because changing a storage client's absent-object semantics should force a review of deletion behavior. That small bit of friction has saved more debugging time than another dashboard panel.
I also keep a deletion event with the record id, object key hash, policy version, attempt count, and outcome. Hashing the key in the event reduces accidental exposure while preserving correlation. Do not log the image URL, signed download URL, or raw identity metadata.
The retry policy should be narrow. Retry network timeouts and rate limits with capped exponential backoff. Do not retry a malformed key forever. After a bounded number of attempts, move the row to failed, retain the reason, and page an operator without putting the photo back into an upload queue.
| Decision | Upload-time action | Scheduled-worker action |
|---|---|---|
| Retention window | Calculate and persist expiresAt
|
Query rows whose deadline has passed |
| Object safety | Generate a unique key | Delete that exact key, never a prefix |
| Retry | Return a clear upload result | Claim, delete, and record an outcome |
| Evidence | Store policy version | Keep a tombstone and audit event |
What should be tested before a retention job reaches production?
Test the policy boundary before testing throughput. A photo expiring at 2026-04-10T12:00:00Z must remain active at 11:59:59Z and be claimable at exactly noon. Add a test for a tenant with two photos sharing a filename but not a generation key. That catches the most expensive class of accidental deletion.
Then test interruptions. Make deleteObject succeed while markDeleted fails, restart the worker, and verify that the second run converges to deleted. Make deleteObject report an absent object and verify the same convergence. Run two workers against the same rows and assert that one claim wins.
For a practical integration test, put three tiny JPEG fixtures in a disposable bucket: one still inside its window, one expired, and one with a past failed attempt. Freeze the clock, run one batch, inspect the object listing, then inspect the tombstones. The fixture need not contain real identity data. It should still exercise MIME validation and the image format assumptions documented by MDN.
I benchmark the boring parts. Measure rows claimed per batch, remote delete latency, and the p95 age of an expired row. I am not sure a larger batch is better on your storage backend; your mileage may vary. A batch of 100 is a starting point, not a promise. Increase it only when the lag metric improves without causing rate-limit retries.
What I would change at scale
At modest volume, one scheduled process and a database lease are enough. At higher volume, partition claims by tenant hash or expiry hour and run bounded workers. Keep the policy service separate from the deletion executor so a policy change cannot mutate already-issued deadlines.
The trade-off is operational surface area. A queue gives smoother load and clearer backpressure, but it adds a broker, delivery semantics, and another dashboard. A database scan is easier to debug and often wins until the expired-row index becomes hot. Choose the smallest system that can demonstrate its deletion evidence to a reviewer.
This approach is not suitable when a regulator requires cryptographic, hardware-backed destruction evidence for every storage replica; use the provider's certified erasure workflow and preserve its attestation. It is also a poor fit for files that must be retained for an active legal hold. In that case, mark the hold explicitly and make the worker skip it with an auditable reason. Do not hide the exception in a cron filter.
The final rule is unglamorous: a retention window is a timestamp, a deletion is a state machine, and a successful run is an observable fact. Keep those three concepts separate and the Node.js code stays small enough to review.
Top comments (1)
Hey, I'm building a small open-source CLI that analyzes a codebase and generates architecture/structure documentation. I'm looking for a few developers willing to run it against a real project and tell me where it gets things wrong.
You don't need to upload your code anywhere just run:
npx @autodocify/autodocs analyze .
Requires Node 20+.
If you try it, I'd especially like to know what it missed or misunderstood.