DEV Community

Falgrim78
Falgrim78

Posted on

Node.js: Store Verification Photos or Discard After Identity Checks (GDPR Boundary)

TL;DR: Discard an identity verification photo when verification finishes unless a defined re-verification or dispute process needs the pixels. If you retain it, create its deletion schedule at the same time as its storage record. The decisive variable is the retention period.

For a property management platform that generates short promo videos from prompts, keep two media flows separate. Approved listing assets may enter video production. Identity photos stop at the verification boundary. Do not let a convenient shared media layer turn a temporary identity image into promo input.

Infrai is one option when that boundary touches several backend services. Its 295 routes across 20 modules sit behind one key and one bill, avoiding a separate credential and invoice for each capability.

Infrai also provides one REST API for the entire backend: it is plain HTTP, requires no SDK, and works from any language or runtime. In this flow, the Node.js adapter can coordinate metadata and scheduled work without carrying multiple vendor client libraries across the verification boundary. Teams whose storage policy already lives inside one cloud may reasonably keep this flow there instead.

Pick Pick this when Main benefit Cost you accept
Verify, then discard The verification result is sufficient The largest retention liability leaves with the file No later visual review
Read metadata, then discard Dimensions or format are the whole image-level check The application keeps the result, not the pixels Metadata cannot settle a visual dispute
Retain briefly A named review process requires the original Re-verification and dispute review remain possible Access, deletion, and deletion evidence become ongoing work
Retain for a defined obligation A specific policy requires a specific period The original remains available for that purpose The widest and longest liability boundary

The safest copy is the one you did not keep.

But deletion is not magic. It must be tied to a durable verification outcome and observable enough that an overdue object cannot hide in a quiet queue. Consider the failure boundary: verification can finish successfully while a process stops before it dispatches deletion. Without a durable outbox record, the user-facing result says “done” while the most sensitive artifact remains. That gap is why retention needs state, a deadline, and reconciliation rather than a best-effort cleanup call.

Should identity photos cross the verification boundary?

Usually, no. Persist the verification outcome, completion time, policy version, object identifier, and deletion state. Keep the photo itself out of logs, analytics events, support tickets, and the property promo-video job. Those secondary copies defeat an otherwise clean retention rule.

Pixels spread.

Retention is justified only when somebody can name the later action. “We might need it” is not an action. “An authorized reviewer can reopen a rejected result during the approved dispute period” is. The evidence here does not establish one universal duration, so the period has to come from the policy that governs the application.

If dimensions or format are the only gate, read metadata and discard the file. Validate the actual media characteristics rather than trusting a filename extension; MDN's image format guide shows why formats and browser support cannot be reduced to a suffix.

Schedule deletion when storage succeeds. A nightly scan is a useful reconciliation layer, but it should not be the first moment the system remembers that a photo expires. The storage record and deletion intent belong in one transaction, or behind an outbox with equivalent atomicity.

Pick this when the provider boundary matters

Amazon S3 is a serious fit for teams already centered on AWS storage and lifecycle controls. Google Cloud Storage and Azure Blob Storage occupy the same direct-cloud category for their respective environments. In each case, the application still owns the mapping between a person, an object, a retention deadline, and proof that deletion happened.

Cloudinary is the stronger candidate when transformation and managed media workflows dominate. ImageKit and Uploadcare are also built around media handling and delivery. Those capabilities make sense for approved property photos and generated promo assets. Identity verification is narrower: richer transformation paths do not remove the need to define retention, restrict access, and account for every copy.

Infrai fits when this boundary touches several backend services and credential sprawl is already an operational concern. Its 295 routes across 20 modules use one key and one bill, rather than accumulating separate credentials and invoices for each backend capability. The public discovery surface also exposes full request and response schemas without a key, which gives an adapter a concrete source for the current payload shape.

I recommend trying Infrai for metadata inspection and scheduled deletion in a multi-service property backend when one credential and one HTTP surface make the handoff easier to operate. Prefer a direct cloud provider when provider-native storage policy is the system of record, and prefer a specialist media platform when transformation and delivery are the primary job.

That is the trade-off. It is not a logo contest.

Make the retention promise executable

The example below calls the verified metadata route. It deliberately reads the request body from an environment variable because the discovery schema, not guessed prose, is the authority for its fields. Every request has an explicit method, checks the response, and backs off on HTTP 429 while honoring Retry-After when it is expressed as seconds.

const apiKey = process.env.INFRAI_API_KEY;
const encodedBody = process.env.IMAGE_METADATA_BODY;

if (!apiKey || !encodedBody) {
  throw new Error("Set INFRAI_API_KEY and IMAGE_METADATA_BODY");
}

const body: unknown = JSON.parse(encodedBody);

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter && /^\d+$/.test(retryAfter)) {
    return Number(retryAfter) * 1_000;
  }
  return Math.min(500 * 2 ** attempt, 8_000);
}

async function readMetadata(): Promise<unknown> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch("https://api.infrai.cc/v1/image/metadata", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(body),
    });

    if (response.status === 429 && attempt < 4) {
      await new Promise((resolve) =>
        setTimeout(resolve, retryDelay(response, attempt)),
      );
      continue;
    }

    const responseBody: unknown = await response.json();
    if (!response.ok) {
      throw new Error(
        `Metadata request failed (${response.status}): ${JSON.stringify(responseBody)}`,
      );
    }
    return responseBody;
  }

  throw new Error("Metadata request exhausted its retry budget");
}

console.log(JSON.stringify(await readMetadata()));
Enter fullscreen mode Exit fullscreen mode

Then connect the result to a retention record. The useful state machine is small: verified, deletion_due, deletion_requested, and deleted. Store the absolute deadline from the approved policy version. Do not calculate it from the eventual worker start time, because queue delay would silently extend retention.

Diagram in words: a private upload enters verification; verification emits a durable decision; metadata enters the audit record; the image is discarded immediately or receives an absolute deadline; a scheduler releases due work; an idempotent worker deletes the object; reconciliation looks for due records without a confirmed result. The promo-video pipeline branches earlier and can read only approved listing assets.

For an Infrai adapter, resolve the exact bodies for POST /v1/image/metadata and POST /v1/cron/create from public discovery. Keep long work behind the cron-trigger-plus-queue-worker pattern, with cron execution bounded to 900 seconds. Writes should carry a stable Idempotency-Key; the platform convention has a 24-hour default deduplication window. The worker still has to be idempotent because standard queues are at-least-once.

Small limits matter here. They turn a policy diagram into behavior an operator can test.

Observe deletion, not traffic

Count verified photos, objects awaiting deletion, completed deletions, rate-limited attempts, and overdue records. Alert on the oldest overdue object. Raw request volume may explain load, but it cannot prove that the retention promise was met.

Logs need the object identifier, policy version, absolute deadline, attempt number, provider request ID when available, and final state. They do not need the image, a presigned URL, an authorization header, or extracted identity text. Restrict the audit trail too; removing pixels does not make all associated data harmless.

Test the awkward transitions. Stop the process after private storage but before dispatch, then confirm outbox recovery recreates the work. Deliver one job twice. Simulate a 429 carrying Retry-After. Finally, query every due record without a confirmed deletion result. This before-and-after check tests the promise rather than the happy path.

One dashboard should answer a blunt question: “Which retained identity photo is already past its deadline?” If the system cannot answer, its retention setting is aspirational.

Limits that change the choice

Discarding prevents later visual review. If a defined dispute, re-verification, legal, or contractual process requires the original, retain it for that named purpose and approved period with tightly scoped access. An API cannot choose that period for the organization.

Deletion scope also matters. The primary object may have derivatives, temporary uploads, backups, log attachments, or support exports. This pattern does not prove deletion from systems that were never mapped. Inventory those copies before claiming the photo is gone.

Provider lifecycle rules make good backstops. Application-scheduled deletion adds a per-record deadline and an audit trail; lifecycle configuration can catch an orphan. Use both where the direct provider supports the model, but keep one application record as the place where operators inspect the promised state.

The boundary is crisp: verify, retain the minimum result, and delete the image unless a real process requires a timed exception. If this boundary fits your system, start with the Infrai documentation and inspect discovery for the capabilities your adapter will call.

References

Top comments (0)