To store and expire generated promo videos safely, make retention part of the API record instead of relying on a storage folder and a cleanup timer. The operational constraint is knowing which uploaded game screenshots, OCR results, intermediate renders, and final videos still have a reason to exist. Process OCR at upload when campaign text must be reviewed before rendering. Process it on demand when most uploads may never become trailers. In both cases, give every asset an explicit expiry and make deletion observable.
TL;DR: model retention as data, not a background-job assumption. A Node.js service should record an asset's purpose, state, expiresAt, and deletion evidence in durable metadata. Object storage holds bytes. A sweeper requests deletion only after the retention rule allows it, then verifies absence before marking the asset deleted. Track age, backlog, and deletion outcomes so storage growth becomes a measurable lifecycle problem rather than a surprise invoice.
From a folder of files to a lifecycle you can see
Picture the pipeline in words. A player uploads a match screenshot. OCR extracts a team name and score. A renderer turns approved text and selected gameplay into a short promotional clip. The application returns the clip to a campaign editor. Later, policy expires the source image, OCR artifact, render fragments, and final video on different schedules.
The weak mental model is a bucket with a nightly cleanup script. The useful mental model is a small state machine with evidence at every transition:
uploaded -> text_extracted -> approved -> rendering -> ready -> deletion_due -> delete_requested -> deleted
Each transition answers an operational question. Is OCR falling behind? Are renders abandoned after approval? Is the deletion worker failing, or is verification merely delayed? A filename cannot answer those questions. Metadata can.
Prove it.
There is also a privacy boundary hiding here. OCR output can preserve words that were visible in a screenshot even after the source image is gone. Treat extracted text as its own retained object, with its own expiry, rather than harmless metadata that lives forever. The same rule applies to thumbnails and render manifests.
A practical record needs fewer fields than many teams expect. Keep a stable asset ID, campaign ID, kind, lifecycle state, object key, byte count, creation time, expiry time, and the last deletion attempt. Do not infer expiry from a path prefix or an object's last-modified timestamp. Those values describe storage layout and mutation, not business intent.
How should an API store and expire generated promo videos?
Start with the OCR timing decision because it determines when derived data appears, how long users wait, and which objects can become unused.
| Decision | Upload-time OCR | On-demand OCR |
|---|---|---|
| Best fit | Text review is mandatory or extraction is reused | Many uploads never become promo videos |
| User-visible wait | Before review | During the first trailer request |
| Main waste risk | Extracting text that is never used | Repeated extraction without request coordination |
| Retention signal | Upload starts the derived asset's lifecycle | First generation request starts it |
Run OCR at upload when a human must validate names, scores, or calls to action before any render starts. The wait is visible early, bad source images can be rejected while the uploader still has context, and the extracted text is ready for repeated trailer variants. The trade-off is immediate compute and another retained artifact for every upload, including screenshots that never reach a campaign.
Run OCR on demand when uploads are speculative and only a small portion become videos. That avoids work on unused images and shortens the lifetime of unused extracted text. It also moves OCR latency and failure into the trailer request path, where users are already waiting for generation. Concurrent requests need coordination so they do not extract the same image twice.
The decision rule is concrete: choose upload-time processing when review is mandatory or reuse is common; choose on-demand processing when conversion from upload to render is low and request latency can absorb extraction. Neither choice changes the retention design. Both should produce the same lifecycle events and attach an expiry to every derived artifact.
That is the trade-off.
Do not optimize this decision from object count alone. Instrument four timings: upload-to-OCR-start, OCR duration, approval-to-render-start, and ready-to-first-download. Add counters for OCR attempts, render attempts, and assets that expire without ever being read. Those signals expose where work is wasted and where users actually wait.
A copyable Node.js retention boundary
Keep policy separate from storage operations. This example uses generic interfaces, injects the clock for deterministic tests, and refuses to equate a deletion request with confirmed absence. That distinction matters. A successful request is an action; verification is evidence.
interface AssetRecord {
id: string;
kind: "source_image" | "ocr_text" | "render_part" | "promo_video";
objectKey: string;
expiresAt: Date;
}
interface AssetRepository {
findDue(now: Date, limit: number): Promise<AssetRecord[]>;
markDeleteRequested(id: string, attemptedAt: Date): Promise<void>;
markDeleted(id: string, verifiedAt: Date): Promise<void>;
recordDeleteFailure(id: string, attemptedAt: Date, reason: string): Promise<void>;
}
interface ObjectStore {
delete(key: string): Promise<void>;
exists(key: string): Promise<boolean>;
}
interface Metrics {
increment(name: string, labels: Record<string, string>): void;
observe(name: string, value: number, labels: Record<string, string>): void;
}
export async function sweepExpiredAssets(
repo: AssetRepository,
store: ObjectStore,
metrics: Metrics,
now: Date,
limit = 100,
): Promise<void> {
const assets = await repo.findDue(now, limit);
for (const asset of assets) {
const labels = { kind: asset.kind };
const lagSeconds = Math.max(0, (now.getTime() - asset.expiresAt.getTime()) / 1000);
metrics.observe("asset_expiry_lag_seconds", lagSeconds, labels);
try {
await repo.markDeleteRequested(asset.id, now);
await store.delete(asset.objectKey);
if (await store.exists(asset.objectKey)) {
throw new Error("object_still_present");
}
await repo.markDeleted(asset.id, now);
metrics.increment("asset_deletions_total", { ...labels, outcome: "verified" });
} catch (error) {
const reason = error instanceof Error ? error.message : "unknown";
await repo.recordDeleteFailure(asset.id, now, reason);
metrics.increment("asset_deletions_total", { ...labels, outcome: "failed" });
}
}
}
The repository must claim due rows so two workers cannot process the same record concurrently. The exact locking mechanism belongs to the database layer; the contract above keeps it out of storage code. Retries should be safe, and deleted should mean verification succeeded, not merely that a request was sent.
Test the boundary with a fixed clock. Include assets expiring exactly at now, assets one millisecond in the future, repeated deletion attempts, and a store that reports the object still present. Also test each asset kind. A final clip and an OCR text object may share machinery while following different retention rules. The batch limit of 100 in the snippet is a starting constraint, not a universal optimum: raise it only after watching sweep duration, repository contention, and the age of the oldest due record together, because a faster-looking worker that starves application queries has moved the problem rather than solved it.
Alerts should describe user and policy risk
One deletion failure is a retry, not an emergency. Alert on sustained symptoms: the oldest due asset keeps aging, the due backlog grows across multiple sweeps, or verified deletions stop while new assets continue to arrive. Page only when policy compliance or service capacity is at risk. A transient error belongs in metrics and logs.
Use low-cardinality metric labels such as asset kind and outcome. Put asset ID, campaign ID, object key, attempt number, and error reason in structured logs. IDs make individual failures traceable, but using them as metric labels creates a distinct time series for each asset.
Storage cost control then becomes arithmetic you can inspect. Measure active bytes by asset kind, bytes already past expiry, creation bytes per hour, and verified deletion bytes per hour. Forecasting does not require a fragile unit-price claim: if creation stays above deletion for disposable objects, retained bytes will rise. Fix lifecycle throughput or policy before debating providers.
Keep a compact dashboard: active bytes, overdue bytes, oldest expiry lag, due records, deletion outcomes, and unused assets expired. Split it by source_image, ocr_text, render_part, and promo_video. That view connects the processing decision to its consequence. Upload-time OCR may show many unread text artifacts; on-demand OCR may instead show longer request-path latency.
What about legal holds and campaign extensions?
An expiry is policy input, not permission to ignore exceptions. A campaign extension should update the relevant expiry before the worker claims the asset. A legal or moderation hold should be an explicit field checked atomically with the due-state transition. Log who or what applied the hold and when, without placing sensitive OCR text in the log message.
Race conditions deserve a deliberate rule. If an editor extends a campaign while deletion is already claimed, either reject the extension with a clear status or create a new render from retained source material. Do not promise resurrection of an object whose deletion has begun.
No hidden grace period.
The second objection is recovery: what if a deletion was accidental? Retention and recovery pull in opposite directions. Define any recovery window as a separate lifecycle state with its own deadline, access controls, and observability. If policy requires immediate erasure, there is no recovery window. The service contract should say which behavior applies rather than letting storage defaults decide.
This architecture leaves provider choice open. More important, it gives developers a crisp operational contract: every generated or derived object has a reason to exist, a deadline, and evidence of removal. Process OCR where the product flow needs it. Measure the consequences. Expire bytes with proof.
Sources
- MDN Web Docs, "Image file type and format guide": https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
Top comments (0)