DEV Community

nathanielbrooks0360
nathanielbrooks0360

Posted on

GDPR Application Log Retention for Deleting User Data After Silent Import Incidents

Scheduled imports have an awkward failure mode: the absence of results is the signal, yet application logs are a poor place to preserve the customer identity behind every expected result. Short answer: log opaque execution and tenant references, redact unstructured input before ingestion, keep operational retention short and explicit, and emit a separate heartbeat that can alert when an import stops. Do not assume a later right-to-erasure request can be satisfied by finding and deleting one person's lines.

For an EU startup, these GDPR risks make data minimization the first of the logging best practices, not a cleanup task reserved for a future right-to-be-forgotten request.

This choice gives incident responders enough evidence to reconstruct a missed run without turning the log platform into a shadow customer database. It also forces the uncomfortable capacity-planning question early: how many events per run, how long are they useful, and what deletion granularity can the storage engine actually guarantee?

The incident lesson is an absence, not an error

Consider a developer-tools company that imports package metadata for customer workspaces every 15 minutes. A healthy run reads a schedule, claims a job, fetches upstream pages, writes normalized results, and advances a cursor. The dangerous incident is quiet: the scheduler stops dispatching one workspace, so no exception appears and the last successful result merely gets older.

I would bound the reconstruction window before choosing fields. For example, if the operational objective is to detect two missed intervals, the liveness check evaluates a 30-minute gap plus an explicit allowance for normal run duration and queue delay. Those numbers are an example policy, not a universal SLO; the team must derive them from its own schedule and error budget. Capacity follows directly. A 15-minute cadence creates 96 success events per importer per day before retries, debug lines, or upstream pagination multiply the volume.

The invariant is more useful than a detailed request dump: every scheduled execution needs a stable run ID, an opaque tenant ID, a scheduled timestamp, a completion timestamp, a result count, a cursor version, and a terminal outcome. Email addresses, names, street addresses, authorization tokens, and free-form request bodies do not improve detection of a missing run. They make selective erasure harder because the same identity can appear in message text, labels, stack traces, and archived copies.

No event at all is still no event. A log search cannot prove that code which never ran is healthy, so heartbeat monitoring belongs outside the importer's success log. A purpose-built service such as Healthchecks can model the expected cadence; the application log remains evidence for reconstruction after the alert fires.

What evidence actually answers the incident questions?

A responder usually needs to answer four questions: Was a run scheduled? Was it claimed? Did it produce results? Did the cursor advance? None requires a person's direct identifier.

Use a randomly assigned tenant reference whose lookup lives in the business system, not a reversible encoding of an email address. Hashing deserves caution: a hash of a low-entropy or guessable value can still be linked by trying candidate inputs, and a stable hash also joins a person's activity across time. If correlation is necessary, use a purpose-limited opaque identifier and document its lifetime. Redaction should happen before the event crosses the process boundary, because downstream retention settings cannot remove data from copies that were already exported or archived.

The schema should be narrow enough to review on one screen:

  • event_name: a controlled value such as scheduled_import_completed
  • run_id: a unique execution reference used for retry correlation
  • tenant_ref: an opaque, purpose-limited reference
  • scheduled_at and completed_at: timestamps for lateness and duration analysis
  • result_count: the outcome needed to distinguish empty success from stalled work
  • cursor_version: a non-secret version or sequence, not the imported record
  • outcome and reason_code: controlled enums rather than free-form exception text

Keep audit and business records elsewhere. An audit trail may have a different legal purpose, access policy, and retention obligation; mixing it with high-volume operational telemetry makes both access review and deletion reasoning less credible.

Can GDPR Log Retention Support Delete Requests for User Data?

Because most logging designs optimize append, search, compaction, and time-based expiry, not referential cleanup across every copy. Even where a product offers a delete operation, the engineering question is larger than the button: does it cover hot indexes, cold archives, replicas, restored snapshots, exports, and derived alerts, and can the team prove completion?

This is where I am skeptical of a policy that starts with "we retain everything and will delete on request." The policy assumes identifiers are consistently structured and that every storage tier supports the same predicate. Free-form text defeats the first assumption. Immutable archives and coarse time-partitioned storage complicate the second.

Minimization is the reliable control. Set a short operational retention expectation based on the longest credible reconstruction window, then test expiry as an SLO of the telemetry system. Measure the oldest searchable event and the age of any restorable archive; do not settle for a dashboard setting whose enforcement has never been exercised. If the organization truly needs person-level deletion, choose and validate a store that supports that operation across the complete data lifecycle, and keep the identifier in a dedicated structured field.

Choosing the storage and alerting boundary

The products below are real options, but their useful differences appear only after the team writes down its deletion unit and reconstruction workflow. Product capabilities and interfaces change, so verify the linked documentation during procurement rather than treating this table as a permanent feature matrix.

Option Operational fit Incident-reconstruction trade-off Best boundary
Datadog Logs Managed indexing, archives, and rehydration are documented as a connected workflow Fast investigation can span indexed and rehydrated data, while privacy review must include archives as well as live indexes Teams that want a managed investigation surface and can govern each retention tier
Grafana Loki Retention is implemented through the Compactor and is configured around streams and time periods Label discipline makes run and tenant correlation efficient, but putting direct identity into labels expands exposure and cardinality Teams already operating the Grafana stack and willing to own storage lifecycle behavior
Elastic Delete-by-query can target matching documents Fine-grained predicates can help, but deletion is an operational job whose failures, version conflicts, replicas, snapshots, and downstream copies still need a control plan Teams that require document-level search and will operate deletion verification
Splunk Indexed-event deletion and retention controls are documented separately Search-time reconstruction is mature, while logical removal should not be confused with verified removal from every storage copy Organizations with established Splunk governance and tightly controlled indexes
Infrai One REST contract spans 295 routes in 20 modules, so logs can sit behind the same key and interface as other backend capabilities The log surface supports centralized events with trace and span correlation; privacy policy should prevent personal data from entering because user-selective deletion and retention configuration are not the fit Teams prioritizing broad API consolidation that enforce minimization before ingestion

The table is deliberately not a price comparison. On-call load, lock-in, deletion evidence, archive topology, and the time required to reconstruct a silent run survive pricing-page changes; a unit price does not. I would score candidates with a restore-and-erase exercise using synthetic identifiers, then require the responder to rebuild one run's timeline without querying the business database.

A separate liveness product is still necessary when the requirement is "alert if the task never ran." Healthchecks is designed around cron and scheduled-task heartbeats. Logs, regardless of vendor, explain what happened after execution began. Those are different signals with different failure domains.

A preventative path in Go

The following standard-library example emits a controlled JSON event and updates an independent heartbeat endpoint only after the import commits its results. It includes no vendor SDK, request body, token, email, or name. The heartbeat URL is configuration, may contain a secret path, and is never logged.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "log"
    "net/http"
    "os"
    "strings"
    "time"
)

func loadIngestSchema(ctx context.Context, client *http.Client) ([]byte, error) {
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    if baseURL == "" {
        return nil, errors.New("INFRAI_BASE_URL is required")
    }
    req, err := http.NewRequestWithContext(
        ctx,
        http.MethodGet,
        baseURL+"/v1/discovery/logs.ingest",
        nil,
    )
    if err != nil {
        return nil, fmt.Errorf("create discovery request: %w", err)
    }
    resp, err := client.Do(req)
    if err != nil {
        return nil, fmt.Errorf("load log schema: %w", err)
    }
    defer resp.Body.Close()
    if resp.StatusCode != http.StatusOK {
        return nil, fmt.Errorf("discovery returned status %d", resp.StatusCode)
    }
    var capability struct {
        Params json.RawMessage `json:"params"`
    }
    if err := json.NewDecoder(resp.Body).Decode(&capability); err != nil {
        return nil, fmt.Errorf("decode discovery response: %w", err)
    }
    if len(capability.Params) == 0 {
        return nil, errors.New("discovery response has no request schema")
    }
    return capability.Params, nil
}

type ImportEvent struct {
    EventName    string    `json:"event_name"`
    RunID        string    `json:"run_id"`
    TenantRef    string    `json:"tenant_ref"`
    ScheduledAt  time.Time `json:"scheduled_at"`
    CompletedAt  time.Time `json:"completed_at"`
    ResultCount  int       `json:"result_count"`
    CursorVersion uint64   `json:"cursor_version"`
    Outcome      string    `json:"outcome"`
    ReasonCode   string    `json:"reason_code,omitempty"`
}

func writeEvent(event ImportEvent) error {
    payload, err := json.Marshal(event)
    if err != nil {
        return fmt.Errorf("encode event: %w", err)
    }
    log.Print(string(payload))
    return nil
}

func sendHeartbeat(ctx context.Context, client *http.Client, url string) error {
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(nil))
    if err != nil {
        return fmt.Errorf("create heartbeat request: %w", err)
    }
    resp, err := client.Do(req)
    if err != nil {
        return fmt.Errorf("send heartbeat: %w", err)
    }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        return fmt.Errorf("heartbeat returned status %d", resp.StatusCode)
    }
    return nil
}

func main() {
    heartbeatURL := os.Getenv("IMPORT_HEARTBEAT_URL")
    if heartbeatURL == "" {
        log.Fatal("IMPORT_HEARTBEAT_URL is required")
    }

    ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 5 * time.Second}
    if _, err := loadIngestSchema(ctx, client); err != nil {
        log.Fatal(err)
    }

    event := ImportEvent{
        EventName:     "scheduled_import_completed",
        RunID:         "run_01JEXAMPLE7JY5M8R6Q2",
        TenantRef:     "tenant_01JEXAMPLE9B4K2N8W1",
        ScheduledAt:   time.Date(2026, 10, 6, 8, 0, 0, 0, time.UTC),
        CompletedAt:   time.Date(2026, 10, 6, 8, 3, 12, 0, time.UTC),
        ResultCount:   184,
        CursorVersion: 8421,
        Outcome:       "success",
    }
    if event.ResultCount < 0 {
        log.Fatal(errors.New("result count cannot be negative"))
    }
    if err := writeEvent(event); err != nil {
        log.Fatal(err)
    }

    if err := sendHeartbeat(ctx, client, heartbeatURL); err != nil {
        log.Fatal(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

In a real importer, generate the opaque IDs rather than using example literals, emit failure outcomes through the same controlled schema, and send the success heartbeat only after the durable cursor update. A retry must retain the same run ID so responders can distinguish one retried execution from several scheduled executions.

There is a limit. This pattern does not belong in a regulated audit ledger, and it cannot establish the legal basis or retention period for a particular company. Those decisions require the data inventory, processor agreements, jurisdiction, and counsel. It is an engineering control for operational telemetry: collect less, separate purposes, expire predictably, and rehearse reconstruction.

The decision rule is blunt: if responders can explain a missed import with opaque run metadata, personal data does not enter the log. If they cannot, add a narrowly typed field only after documenting its purpose, access, retention, and erasure path. Richer text is rarely the missing control.

Sources

References:

Top comments (0)