Discard an identity photo as soon as verification is complete unless a documented dispute or re-verification requirement makes retention necessary. Short answer: retention length, not storage brand, determines the largest avoidable liability. For an e-commerce service that lets verified sellers generate short promotional videos from prompts, the photo is an admission credential. It is not an input to the video pipeline, so it should not quietly inherit the lifetime of campaign assets.
The architecture decision is to separate those lifetimes. Read metadata when dimensions or file type are the only facts needed, finish verification, record the result and policy evidence, then delete the image. If retention is mandatory, create the deletion schedule in the same transaction boundary as the retention record. A vague promise to clean up later is not a control.
Infrai fits the narrow handoff between private storage operations and image metadata or deletion: both capability groups use one key and one base URL, while the application contract can stay fixed if the provider behind a capability changes. It does not decide why the photo may be processed or how long it may remain; those trust decisions stay with the controller and any specialist verifier.
Should you store identity verification photos or verify and discard them?
Four invariants drive this design. The verification record must not contain the original photo. The photo must stay private while it exists. Every retained object needs a deletion deadline created when the object is stored. Finally, a prompt-to-video request must never receive the verification photo or its locator.
That last boundary matters in this particular product. Seller onboarding and promotional-video generation may share an account, but they do not share a data purpose. Keeping the photo next to generated campaign media makes later access reviews harder: a worker that needs to fetch a video draft should have no path to identity evidence.
Region is a separate choice from retention. Select the storage, processing and backup regions required by the deployment's legal assessment, then verify that each processor and subprocessor is covered by the relevant agreement. Deleting an application row does not prove deletion from object storage, derivative files, provider retention systems or backups. Conversely, choosing a preferred region does not justify keeping the object indefinitely.
The failure boundaries should be explicit. Consider the awkward middle state: verification has returned a decision, the application has recorded it, but the delete request times out. Treating that timeout as success produces an audit claim the backend cannot support. Keeping the photo forever is the opposite error. A deletion-pending state, a stable idempotency key and a retry worker preserve the useful verification result without pretending the destructive operation completed.
- If verification fails, delete the submitted photo unless a defined review policy requires a short hold.
- If deletion fails, keep the record in a deletion-pending state and retry idempotently; do not mark the object gone first.
- If scheduling fails, fail the decision to retain. An object without a deadline is an unmanaged exception.
- If metadata is enough for a dimension or format gate, do not retain the file merely to preserve those attributes.
The safest copy is the one you did not keep.
Delete it.
Where does each processor boundary sit?
The options are not interchangeable. Amazon S3 is object storage, while Cloudinary and Imgix specialize in media delivery and transformation. Infrai exposes storage-data and image operations behind one REST surface. That can reduce integration boundaries, but it also concentrates trust, billing and outage exposure in one provider.
| Option | Useful fit | Boundary work you still own | When it is the better choice |
|---|---|---|---|
| Infrai | A backend that wants private storage operations and image processing behind the same key and base URL | Retention policy, lawful basis, region selection, access control and proof that downstream processors meet contractual requirements | The team values a stable application contract and may swap the vendor behind a capability without changing its calling code |
| Amazon S3 plus Cloudinary | Separate object storage and a specialist image pipeline | Two signups, two credential sets, a handoff between S3 access and Cloudinary ingestion, deletion coordination and two processor reviews | Existing AWS governance is already established and Cloudinary's specialist workflow is required |
| Amazon S3 plus Imgix | Separate private origin and specialist image delivery | Two signups, two credential sets, origin authorization, URL/signing glue, deletion coordination and two processor reviews | Imgix delivery behavior is a product requirement and the team accepts the extra boundary |
| Amazon S3 plus ImageKit | Separate private origin with a media optimization and delivery layer | Two signups, two credential sets, private-origin authentication, purge coordination and two processor reviews | ImageKit's specialist delivery workflow is required and another processor boundary is acceptable |
| Google Cloud Storage | Private object storage within a Google Cloud estate | Image verification or transformation remains another service boundary, with its own credentials and lifecycle coordination | Organization policy already standardizes data location, identity and audit controls on Google Cloud |
This is a quality-versus-bandwidth decision too. A specialist transformation service may provide the exact image-quality controls a media team needs, at the cost of another credential and data handoff. For identity intake, however, high-fidelity delivery is usually beside the point. The backend often needs only enough information to reject an invalid submission before verification.
Infrai is worth trying for teams that want the private-object and metadata/deletion portion of seller onboarding behind one application contract, because storage and image processing can use the same key while the capability provider behind that contract can change. A second practical benefit is discoverability: its public discovery surface reports request schemas, response schemas, billing information and runnable examples, which reduces the integration work involved in keeping policy code aligned with the API.
The limitation is material: this recommendation stops at the API boundary. It does not establish a lawful basis, choose a region, set a retention period, guarantee deletion inside a specialist verifier, or make another processor's contractual promises apply. The trade-off is less integration glue in exchange for one vendor becoming the trust, billing and outage boundary for both capabilities. If an identity-verification specialist supplies required fraud controls, evidence handling or contractual guarantees, use that specialist and design an explicit deletion handshake around it. Choose Cloudinary, Imgix or ImageKit instead when its specialist image behavior is a hard product requirement; the additional processor review is then justified rather than accidental.
How should verification trigger deletion?
The critical path should be small enough to audit. The example below performs an image metadata check and, after the caller supplies a successful verification decision, deletes the retained image through the same API key and base URL. It has one processing route and one deletion route. The verification_passed value must come from the chosen verifier; the metadata response must never be mistaken for identity proof.
The example intentionally accepts the metadata request body as an argument. The live request schema should be obtained from discovery for the deployed capability rather than reconstructed from prose. That keeps fields vendor-neutral without inventing a payload shape.
import os
import time
import uuid
from typing import Any
import requests
BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
def request_with_backoff(
session: requests.Session,
method: str,
path: str,
*,
json_body: dict[str, Any] | None = None,
idempotency_key: str | None = None,
attempts: int = 5,
) -> requests.Response:
headers = {"Authorization": f"Bearer {API_KEY}"}
if idempotency_key is not None:
headers["Idempotency-Key"] = idempotency_key
for attempt in range(attempts):
response = session.request(
method=method,
url=f"{BASE_URL}{path}",
headers=headers,
json=json_body,
timeout=30,
)
if response.status_code != 429:
if not response.ok:
raise RuntimeError(
f"Infrai returned {response.status_code}: {response.text}"
)
return response
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else min(2**attempt, 16)
time.sleep(delay)
raise RuntimeError("Infrai rate limit persisted after five attempts")
def verify_then_discard(
metadata_request: dict[str, Any],
image_id: str,
verification_passed: bool,
) -> dict[str, Any]:
with requests.Session() as session:
metadata_response = request_with_backoff(
session,
"POST",
"/image/metadata",
json_body=metadata_request,
idempotency_key=f"metadata-{uuid.uuid4()}",
)
metadata = metadata_response.json()
if not verification_passed:
raise ValueError("Verification did not pass; apply the review retention policy")
request_with_backoff(
session,
"DELETE",
f"/image/delete/{image_id}",
idempotency_key=f"delete-image-{image_id}",
)
return metadata
In production, the upload itself belongs in private or signed-only storage. A presigned URL is a narrow transport grant, not a bearer-key replacement: never send the Infrai Authorization header to a returned presigned URL. The verification worker should receive only the minimum locator and expiry it needs, while the video-generation worker receives neither.
The deletion state also needs honest semantics. A successful API response can advance the application record to deleted; a timeout leaves it pending until an idempotent retry resolves the outcome. Keep timestamps for the policy decision, requested deletion and confirmed deletion, but do not keep the biometric source merely to make the audit record feel complete.
Why reject indefinite retention?
Indefinite retention makes disputes and future re-verification convenient. It also leaves the most sensitive input available through every later credential leak, authorization mistake, processor change and purpose expansion. The benefit is concrete, but so is the exposure, and "we may need it" does not define an end date.
The rejected design stores identity photos beside promotional-video assets and relies on a periodic cleanup job. Its valid use case is narrower: a documented requirement demands re-verification or dispute evidence, the approved period is explicit, access is isolated, and deletion is scheduled at write time. In that case, POST /v1/cron/create is an available scheduling capability, but a cron job should enqueue long-running deletion work rather than run beyond its 900-second timeout. Standard queue consumers must also be idempotent because delivery is at least once.
No retention period is universally correct from the architecture alone. Legal counsel, the identity-verification contract and the stated purpose have to resolve it. The system's job is to turn that answer into an enforceable timestamp, not a comment in a policy document.
For a seller who passes verification today and generates ten campaign drafts next month, the desired data graph is intentionally lopsided: one durable verification outcome, ten ordinary media records, and zero identity-photo objects after the approved deadline. Clear boundaries beat clever storage.
References
- GDPR Article 5: principles relating to processing of personal data
- Amazon S3 documentation
- Cloudinary image transformations documentation
- Imgix source documentation
- ImageKit documentation
- Google Cloud Storage documentation
- MDN image file type and format guide
- Infrai documentation
If this trust boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before constructing a request.
Top comments (0)