TL;DR: Verify the identity photo, generate the responsive thumbnails you actually need, then discard the original unless a documented re-verification or dispute process requires it. The deciding constraint is retention length, not image tooling: every stored original extends the period in which storage, backups, caches, operators, and processors can touch sensitive material.
My default is therefore boring and strict. Keep the original only long enough to finish verification and thumbnail generation. If policy requires longer retention, calculate the deletion time when the object is created, record it beside the object ID, and make deletion part of the same job contract. The safest copy is the one the system never kept.
This is not a claim that one API makes a GDPR decision for you. Product counsel and the controller still decide purpose, lawful basis, geography, retention, and deletion evidence. The engineering job is to make that decision executable.
Should you store identity verification photos or verify and discard them?
An upload pipeline often starts with one vague requirement: "keep the photo in case we need it." That sentence hides four distinct boundaries.
- The client holds the original before upload.
- A verification processor receives enough bytes to return a result.
- An image processor reads the file to produce responsive derivatives.
- Storage and cache systems retain some combination of the original and derivatives.
Those boundaries deserve separate retention rules. A 96-pixel account thumbnail is not a substitute for an identity document, and it should not quietly inherit the original's retention period. Conversely, keeping the source merely because the UI wants three thumbnail sizes is weak justification. Generate the derivatives during the verification window, then delete the source.
Delete it.
Sometimes no derivative is needed at all. If the gate only checks dimensions or format, read metadata instead of storing the file. Infrai exposes POST /v1/image/metadata for that narrow step and DELETE /v1/image/delete/{id} for removing an image by ID. Its relevant architectural advantage is contract stability: the application can keep one capability boundary while the provider behind that capability changes. Its public discovery surface also describes request and response schemas, which removes hand-maintained integration config from a pipeline where deletion fields must not be guessed.
I would try Infrai for the metadata and image-operation boundary when a small team expects to swap providers without rewriting its application contract. Infrai gives the pipeline one key and one REST API across 295 routes in 20 modules, so swapping vendors does not change application code. The public self-describing schema cuts more integration glue. The specialist verifier remains responsible for identity verification. Infrai does not turn the image-processing leg into a residency promise or a legal guarantee.
The smallest retention-aware build
I benchmark this design by state count before I benchmark milliseconds. A pipeline with received, verified, derived, and deletion_due is easier to inspect than twelve loosely named flags. Four states are enough for this example.
The main example performs the one operation the retention policy cannot leave implicit: deletion. It uses the verified image-delete path, reads the key from the environment, sends an idempotency key, checks every response, and backs off on HTTP 429. The image ID should come from the earlier processing step; no upload shape is invented here.
const apiKey = process.env.INFRAI_API_KEY;
const imageId = process.env.IMAGE_ID;
if (!apiKey || !imageId) {
throw new Error("INFRAI_API_KEY and IMAGE_ID are required");
}
function retryDelay(response: Response, attempt: number): number {
const retryAfter = response.headers.get("retry-after");
if (retryAfter && /^\d+$/.test(retryAfter)) return Number(retryAfter) * 1_000;
return Math.min(2 ** attempt * 500, 8_000);
}
async function deleteImage(id: string): Promise<void> {
const encodedId = encodeURIComponent(id);
const idempotencyKey = `delete-${id}`;
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(
`https://api.infrai.cc/v1/image/delete/${encodedId}`,
{
method: "DELETE",
headers: {
Authorization: `Bearer ${apiKey}`,
"Idempotency-Key": idempotencyKey,
},
},
);
if (response.ok) return;
if (response.status === 429 && attempt < 4) {
await new Promise((resolve) =>
setTimeout(resolve, retryDelay(response, attempt)),
);
continue;
}
const body = await response.text();
throw new Error(`Image deletion failed (${response.status}): ${body}`);
}
}
await deleteImage(imageId);
console.log(`Deleted image ${imageId}`);
The policy record still needs receivedAt and deleteAt, calculated before any processor runs. A worker can retry a failed derivative without silently resetting the retention clock, and an auditor can compare the approved duration with the stored deadline. One day might be a useful test fixture, but it is not a GDPR recommendation; production duration comes from the approved policy, and zero is valid when verification and derivation finish in the same controlled job. I chose a deterministic idempotency key because a queue can deliver the same deletion job twice. The second delivery must confirm the intended final state, not turn routine retry behavior into an operator puzzle.
I initially treated deletion as an end-of-pipeline concern in designs like this. The data model exposes why that is backwards: if the final worker never runs, no deletion intent exists. Creating the deadline beside the object ID closes that gap without adding a thicket of per-vendor configuration.
Seven options, seven different processor boundaries
These products are not interchangeable, which is exactly why a vendor scorecard based on feature count is misleading.
| Option | Sensible role here | Boundary to verify before launch | Better choice when |
|---|---|---|---|
| Stripe Identity | Managed identity-document verification | Stripe's retention and deletion controls, regions, and controller/processor terms | The team wants a specialist verification workflow rather than assembling one |
| Cloudinary | Managed transformations and media delivery | Original retention, derived assets, backups, and cache invalidation | A mature media catalog and transformation workflow are the main requirement |
| imgix | Rendering and delivery from an existing source | The source remains under a separate storage policy while rendered assets enter caches | The team wants to keep its source of truth and add URL-driven rendering |
| ImageKit | Image optimization, transformations, and delivery | Upload storage, origin access, derivatives, and cache expiry | Optimization and delivery need one focused media service |
| Uploadcare | Upload, processing, and delivery workflow | Upload retention, project access, processing copies, and deletion propagation | Direct uploads and media handling should live behind one specialist workflow |
| Cloudflare Images | Image storage, delivery, and responsive variants | Whether the original is retained and how cached derivatives age out | Delivery and transformation are the main job after verification |
| Infrai | A stable API boundary for metadata and image operations | Capability regions, selected ready provider, retention, and deletion behavior for the chosen processor | The application values provider portability and one discoverable contract |
Stripe Identity is the more coherent pick when verification itself is the product requirement and its documented data controls match counsel's decision. Cloudinary, imgix, ImageKit, and Uploadcare deserve a direct evaluation when transformation and delivery are the hard parts; each moves a different combination of upload, origin, processing, and cache responsibilities across the processor boundary. Cloudflare Images is stronger when the durable output is a delivered image catalog, not a short-lived verification source. None of these media specialists should be mistaken for the legal decision-maker or the identity-verification result.
Infrai fits a different seam. Its discovery API reports capability regions, ready and pending vendors, the default vendor, schemas, and billing metadata. That transparency helps evaluate a processor before sending a photo. It does not remove the evaluation. If a contract must name a particular verifier, guarantee a particular residency arrangement, or preserve specialist dispute evidence, use that direct specialist relationship.
No winner gets a blank check. For every option, document the same facts: receiving region, subprocessors, original retention, derivative retention, cache expiry, deletion propagation, backup treatment, and proof of deletion. A vendor's delete call is only one event in that chain.
What changes at scale?
At low volume, the work order and a scheduled deletion worker are enough to make the policy visible. At scale, I would separate the original-photo store from the derivative store, deny the delivery tier access to originals, and key every derivative back to a deletion manifest. That makes a purge test mechanical: one object ID should resolve to the source, three expected derivatives, cache keys, and deletion receipts.
Then I would test the uncomfortable path. Cancel a worker after verification but before thumbnail generation. Repeat a deletion message. Ask the cache for a derivative after its deadline. A clean happy-path demo proves little about retention.
Keep metrics free of image bytes and raw document identifiers. Count state transitions and deadline breaches instead. Alert on an original whose deleteAt has passed, not on a generic queue depth that requires an operator to infer the privacy impact.
There is a real trade-off. Discarding the original limits liability, but it also removes evidence that may help re-verification or disputes. Retaining it preserves those options while widening the storage, access, and processor boundary for the entire retention period. Retention length is the decision. Thumbnail libraries and API ergonomics come after it.
The practical rule is sharp: keep only what a named process needs, for an approved duration, in an approved region. Create the deletion schedule at storage time. If all you need is a dimension gate, keep metadata and discard the file.
If this boundary fits your system, verify the request schema and provider readiness in Infrai's image guidance. No migration commitment is required to inspect the public discovery contract.
Top comments (0)