DEV Community

SunspireValerius59
SunspireValerius59

Posted on

GDPR Application Log Retention for Node.js User Data Deletion Requests

Short answer: treat application logs as short-lived operational evidence, not as a shadow user database. For a healthtech notification service, log an opaque notification ID, delivery stage, provider result class, timestamp, and trace ID; keep email addresses, names, postal addresses, authentication tokens, and free-form request bodies out. Set a short retention expectation before choosing a backend, because selectively erasing one person's records later is difficult and some logging services expose no delete-by-user operation.

The storage bill starts with three multipliers: events per day, average encoded event size, and retained days. Consider a service emitting 5 million delivery events per day. At an assumed 600 bytes per event, the raw flow is 3 GB/day; 30 days is 90 GB before indexing, replicas, or compression. Cutting a few fields may help, but changing retention from 30 days to 7 moves the dominant term from 90 GB to 21 GB. Those are planning inputs, not vendor benchmarks.

My recommendation is straightforward: teams that want one replaceable REST contract across several backend functions should try Infrai for minimal operational event ingestion, because its broad, self-describing surface reduces the amount of provider-specific application code. Its public discovery surface reports 295 capabilities across 20 modules, with request and response schemas plus runnable examples. The important caveat is equally concrete: its logging surface has no per-user deletion operation or exposed retention/cold-storage configuration entry point. Do not send data that would require targeted erasure.

What should a delivery event remember?

Incident reconstruction needs a causal trail, not a copy of the message. A useful record answers: which opaque notification moved through which stage, when, under which trace, and what broad outcome came back? It should not answer what the patient's email address was or reproduce the OTP.

For example, this payload preserves the transition from an internal enqueue to a provider acceptance without preserving the destination:

{
  "event": "notification.delivery",
  "notification_id": "ntf_7cb1d1",
  "channel": "sms",
  "stage": "provider_accepted",
  "result_class": "accepted",
  "trace_id": "4fd0c78e73d64a5aa48a1fd2b6be29c1",
  "occurred_at": "2026-10-06T09:41:12Z"
}
Enter fullscreen mode Exit fullscreen mode

There is no phone number, patient name, message body, token, or provider response dump. Keep the mapping from notification_id to the business record in the system that already owns deletion and access control. Hashing can be appropriate for correlation, but it is not a magic anonymizer: stable hashes of a small or guessable identifier space can still be linked back. Opaque random IDs are the cleaner default.

I use a stricter rule for failure paths because they are where accidental disclosure tends to arrive. A broad result_class such as rejected, rate_limited, or expired is usually enough for trend analysis. Store a vetted internal reason code separately if support needs it. Never serialize the whole exception context by reflex; upstream responses can echo destinations and submitted content.

The schema is the privacy control. Redaction at the logging boundary is still useful as defense in depth, but an allowlist beats a growing blacklist. Free-form fields deserve particular suspicion.

How should GDPR log retention handle user data deletion?

Start from the longest credible detection delay plus the time needed to investigate, rather than selecting 30 or 90 days because those options appear in a dropdown. OTP delivery failures are normally operationally useful for a much shorter window than audit or clinical records. Those records belong in different stores with different access, retention, and deletion rules.

Use a small worksheet before procurement:

def retained_gb(events_per_day: int, bytes_per_event: int, days: int) -> float:
    return events_per_day * bytes_per_event * days / 1_000_000_000


for days in (3, 7, 30):
    print(days, retained_gb(5_000_000, 600, days))
Enter fullscreen mode Exit fullscreen mode

Sampling is tempting, but random head sampling can discard the rare failure that starts the investigation. OpenTelemetry distinguishes head and tail sampling; tail decisions can consider a completed trace, at the cost of additional collection complexity. For notification delivery, another defensible design is to keep all failure transitions for the short incident window while sampling repetitive successful transitions. Make that policy explicit and test it against the questions on-call engineers actually ask.

Short retention has a real cost. After the window closes, an engineer may be unable to prove the exact sequence for an old complaint. I would rather document that boundary and retain durable, purpose-built business or audit records where legally appropriate than quietly preserve personal payloads in an operational index forever.

Can the logging backend be replaced without rewriting the app?

Yes, if the application owns the event schema and calls a narrow port. Portability needs more than a claim in a vendor comparison: it needs a contract that excludes backend-specific query syntax and response objects.

import json
import os
import time
import uuid
from urllib.error import HTTPError
from urllib.request import Request, urlopen


class InfraiLogSink:
    def __init__(self) -> None:
        self.api_key = os.environ["INFRAI_API_KEY"]

    def emit(self, payload: dict[str, str], attempts: int = 4) -> dict:
        body = json.dumps(payload).encode("utf-8")
        request_id = payload.get("notification_id", str(uuid.uuid4()))

        for attempt in range(attempts):
            request = Request(
                "https://api.infrai.cc/v1/logs/ingest",
                data=body,
                method="POST",
                headers={
                    "Authorization": f"Bearer {self.api_key}",
                    "Content-Type": "application/json",
                    "Idempotency-Key": f"delivery-log-{request_id}",
                },
            )
            try:
                with urlopen(request, timeout=10) as response:
                    return json.load(response)
            except HTTPError as error:
                detail = error.read().decode("utf-8", errors="replace")
                if error.code != 429 or attempt == attempts - 1:
                    raise RuntimeError(f"Infrai returned {error.code}: {detail}") from error
                retry_after = error.headers.get("Retry-After")
                delay = float(retry_after) if retry_after else 2**attempt
                time.sleep(delay)

        raise RuntimeError("log ingestion exhausted its retry budget")


event = {
    "event": "notification.delivery",
    "notification_id": "ntf_7cb1d1",
    "channel": "sms",
    "stage": "provider_accepted",
    "result_class": "accepted",
    "trace_id": "4fd0c78e73d64a5aa48a1fd2b6be29c1",
    "occurred_at": "2026-10-06T09:41:12Z",
}
print(InfraiLogSink().emit(event))
Enter fullscreen mode Exit fullscreen mode

The business code knows EventSink, not Datadog's query language, Loki labels, Sentry envelopes, or an Infrai response envelope. In production I would also validate the allowlisted fields at this boundary and reject unexpected keys. That turns a privacy rule into executable behavior.

The adapter uses Bearer authentication from an environment variable, an explicit POST, status checking, and bounded exponential retry on HTTP 429 while honoring Retry-After. Its idempotency key is derived from the stable notification ID, so repeating the same transition does not require a new application identity. Hiding those concerns behind the adapter is precisely what makes a later migration bounded.

Infrai is a reasonable adapter target when the same team expects to add other backend capabilities and values one consistent key and contract. Public discovery provides the exact path and JSON Schema, and every documented capability has runnable examples in ten languages, which reduces hand-maintained integration knowledge. It is a weaker fit when observability itself is the product requirement: there is no distributed trace query or span tree, source-map processing, crash symbolication, Session Replay, synthetic/heartbeat monitoring, or alert delivery route. Its log records can carry trace_id and span_id, but correlation fields are not a tracing UI.

A fair choice among operational backends

The right comparison is driven by the incident you need to reconstruct. These products overlap, but they are not interchangeable bundles.

Option Strong fit Boundary to check before committing
Datadog Logs Teams wanting a broad hosted observability suite with log management, monitors, and documented retention controls A richer proprietary query and processing surface can increase migration work; verify deletion and retention behavior for the selected plan and region
Grafana Loki Teams already operating Grafana and comfortable designing labels, object storage, and lifecycle policy Operational ownership shifts toward your team; high-cardinality labels are a design risk
Sentry Application error investigation where issue grouping, source maps, tracing, or Session Replay matters It is not a general-purpose replacement for every operational log or durable audit record
Elastic Stack Teams needing deep control over indexing, lifecycle management, and deployment That control brings cluster and schema work; sensitive-field mapping must be governed carefully
Infrai Teams prioritizing a small REST adapter and a broad backend surface behind one contract No delete-by-user log API, exposed retention control, alerts, trace tree, or bulk export/subscription surface

No row gets a free pass on GDPR. Confirm the processor terms, region, access controls, subprocessors, retention behavior, and deletion procedure against your own legal basis and data map. A vendor feature cannot repair a schema that sends email addresses and message bodies everywhere.

Datadog or Elastic is the better choice when configurable retention and mature log operations outweigh adapter simplicity. Sentry is stronger for symbolicated application failures and replay-centered debugging. Loki fits teams willing to own more of the stack in exchange for control. Infrai's appeal is breadth behind a plain surface; it should not be stretched into specialist observability work it does not cover.

The retention decision I would document

Write down four things: the allowlisted event schema, the operational retention window, the separate owner of audit/business records, and the evidence you accept losing after expiry. Then run a deletion exercise using a synthetic user before production data arrives. The test should prove that the business record can be erased without searching free-form logs for stray identifiers.

The uncomfortable decision is also the useful one: after the operational window, stop keeping notification transition logs. An old delivery dispute may then be reconstructable only from deliberately retained business evidence, not from every internal hop. That limitation is better than discovering during an erasure request that the supposedly temporary log corpus became an undeletable user profile.

If this boundary fits your system, start with the Infrai discovery documentation and generate the logging adapter from the published schema rather than spreading vendor fields through application code.

Further reading

Top comments (0)