Short answer: For identity verification photos under GDPR, verify and discard by default; store them only when a documented dispute or re-verification need outweighs the added retention liability.
| Decision path | What remains after verification | Best fit | Main trade-off |
|---|---|---|---|
| Verify and discard | Verification result, policy version, timestamps, and operational evidence | A B2B SaaS that needs an access decision, not a photo archive | A later manual review cannot use the original image |
| Store, then delete on schedule | Original and derived images plus the verification record | A workflow with a defined review or dispute window | Every retained copy extends the system that must be governed |
My default decision: discard the source and every compressed or moderated derivative as one operation after the verification result is committed. Retention is the exception. This choice isn't about storage cost; it is about keeping a one-person SaaS small enough to understand, test, and ship weekly.
The table is only a starting point. A real decision needs two answers: does the business process require the image later, and can the deletion promise be proved across every place the bytes traveled? Moderation coverage matters here because rejected uploads, thumbnails, and temporary inputs are still part of the photo lifecycle. Ignoring those branches makes a neat policy diagram and a messy production system.
Should You Store Identity Verification Photos or Verify and Discard Under GDPR?
Choose verify and discard when the durable business object is a decision such as verified, rejected, or manual_review, and no later process requires a person to inspect the original. Keep the decision record. Don't keep the image merely because it might be useful someday. The safest retained copy is the one that doesn't exist.
Choose scheduled retention when a specific workflow must revisit the evidence. A dispute window is a coherent reason. “Analytics” is not yet a reason; it is a label that hides unanswered questions about who will use the file, for what output, and for how long. The retention period should come from that workflow, then be reviewed by privacy counsel against the actual jurisdiction and product role. I'm not sure any generic number can survive contact with every B2B contract, regulator, and identity method.
The catch is straightforward: discard is not suitable when an authorized reviewer must compare the submitted photo with later evidence. In that case, retain for the defined window and make deletion part of the write path. Conversely, storage is a poor default when the application only needs a dimension check, moderation result, or verification outcome. Metadata can preserve what happened without preserving the file itself.
This is a liability comparison, not a claim that one architecture makes legal review unnecessary. The product still needs a plain-language purpose, an owner for the policy, and a response for deletion requests or contractual obligations. Engineering can make those decisions enforceable. It cannot invent them.
The Real Boundary Is the Copy Graph
Teams often debate whether to keep original.jpg and miss the six less obvious copies around it. An upload can pass through a request buffer, temporary object, decoder, resized derivative, moderation input, retry queue, log attachment, and backup. Compression changes the representation, but it does not turn an identity photo into harmless data. A smaller copy is still a copy.
Draw the copy graph before choosing a retention path. Start at the browser upload and follow the bytes through preprocessing, moderation, verification, serving, observability, and deletion. For each node, record the purpose, owner, maximum lifetime, access boundary, and deletion mechanism. One long paragraph in an architecture note is useful here because the awkward branches matter: if a moderation service receives the source while the verifier receives a normalized derivative, a successful request creates two external processing paths; if a rejected image is placed on a manual-review queue, the failure branch may retain data longer than the happy path; and if an application retry serializes the full request body, a transient queue can quietly become another store. The right unit of review is therefore the whole graph, not the primary bucket.
No mystery.
Moderation coverage should be explicit for the same reason. Decide which formats enter the pipeline, which frames or pages are checked, whether the original or normalized image is evaluated, and what happens to rejected input. MDN's image format guide is a useful inventory of common web image formats, but format support is not a retention policy. The operational question is whether every accepted representation follows the same terminal deletion rule.
The durable record can usually be much narrower than the photo. It might contain a subject reference, result, verifier policy version, moderation outcome, timestamps, and a deletion receipt or state transition. Avoid copying image bytes into logs or event payloads. Also avoid recording more metadata than the dispute and audit process can explain. “Keep metadata instead” is a design prompt, not permission to collect an unlimited shadow profile.
Make Deletion Part of the Transaction
A weekly shipping cadence favors one boring interface with two policy modes. The application should not let each feature team improvise storage. In a solo operation, there is no spare hour for reconciling three cleanup scripts after a policy change — outsource undifferentiated image processing behind an interface, but keep lifecycle control in the application.
This TypeScript example uses pseudonymous adapters. The example values are product policy inputs, not legal defaults. It deliberately returns a narrow receipt rather than the image:
type RetentionPolicy =
| { mode: "discard" }
| { mode: "retain"; deleteAt: Date };
type VerificationReceipt = {
subjectRef: string;
result: "verified" | "rejected" | "manual_review";
moderation: "accepted" | "rejected";
policyVersion: string;
verifiedAt: string;
imageState: "deleted" | "scheduled";
};
interface ImagePipeline {
normalize(input: Uint8Array): Promise<Uint8Array>;
moderate(input: Uint8Array): Promise<"accepted" | "rejected">;
verify(input: Uint8Array): Promise<VerificationReceipt["result"]>;
}
interface PrivateObjectStore {
put(input: Uint8Array): Promise<{ objectId: string }>;
delete(objectId: string): Promise<void>;
scheduleDelete(objectId: string, at: Date): Promise<void>;
}
interface ReceiptStore {
commit(receipt: VerificationReceipt): Promise<void>;
}
async function processIdentityPhoto(
subjectRef: string,
source: Uint8Array,
policy: RetentionPolicy,
images: ImagePipeline,
objects: PrivateObjectStore,
receipts: ReceiptStore,
): Promise<VerificationReceipt> {
const normalized = await images.normalize(source);
const moderation = await images.moderate(normalized);
const result = moderation === "accepted"
? await images.verify(normalized)
: "rejected";
const stored = await objects.put(normalized);
const verifiedAt = new Date().toISOString();
const imageState = policy.mode === "discard" ? "deleted" : "scheduled";
const receipt: VerificationReceipt = {
subjectRef,
result,
moderation,
policyVersion: "identity-photo-v3",
verifiedAt,
imageState,
};
await receipts.commit(receipt);
if (policy.mode === "discard") {
await objects.delete(stored.objectId);
} else {
await objects.scheduleDelete(stored.objectId, policy.deleteAt);
}
return receipt;
}
The code exposes an important ordering choice. The receipt is committed before deletion, so the application preserves the verification outcome. A production design also needs idempotency: repeating deletion should converge on “absent,” and repeating a scheduled job should not extend the deadline. Do not silently turn a cleanup failure into indefinite retention. Record deletion state, alert on overdue objects, and retry without placing the photo itself in the alert.
There is a sharper edge in the example too. The normalized image is stored only after verification, which keeps the illustrative interface readable, but a real pipeline may need a private temporary object before calling processors. Add that object to the same lifecycle state machine. Never leave it to a generic cache timeout unless that timeout is the reviewed retention rule.
Test the unhappy paths. Cancel the request after normalization. Reject it during moderation. Repeat the worker. Advance the clock beyond deleteAt. Then assert that originals, derivatives, temporary objects, and queued payloads are absent while the receipt remains. Those tests produce more confidence than a policy sentence because they exercise the places where retention expands by accident.
When the Runner-Up Is the Better Choice
Scheduled retention wins when the business can name the later reviewer, the exact event that opens review, and the event or deadline that closes it. Stick with retention when removing the photo would make a promised appeal or re-verification workflow impossible. Keep access narrow, make the deadline visible beside the object, and test deletion as a release criterion.
Discard wins for a service that only gates account access and can stand behind the recorded result. It reduces the copy graph, backup scope, access-control surface, and operational work. That last part matters to an indie SaaS: revenue per hour improves when the system has fewer exceptional data paths to audit, not when another dashboard is added.
Either choice has a failure mode. Retention can become permanent through vague extensions. Immediate discard can erase evidence that the product explicitly promised to review. The decision rule is simple enough to revisit each quarter: name the post-verification use. If nobody owns that use, remove the bytes. If somebody does, set the deletion time when storage begins and make overdue deletion observable.
Ship the policy with the code.
Top comments (0)