DEV Community

YvesSterling6854
YvesSterling6854

Posted on

After an API Key Leak in 2026 — Establish Actual Touches with Postgres Logs

An API key leak creates two competing operational constraints: stop further access, but preserve enough trustworthy evidence to learn what the key already touched. Revoke too early without capturing the relevant records and responders may lose the easiest correlation handle; wait without containment and the exposure continues.

TL;DR: create a replacement credential, deploy it alongside the still-valid old one, verify the replacement is serving production traffic, then revoke the leaked key. For the investigation, build one UTC-ordered usage series from immutable gateway, application, data-access, and control-plane logs. Join on a keyed fingerprint, request or trace ID, authenticated principal, source context, action, resource, outcome, and timestamp. The blast radius is the set of resources and operations supported by those records, not every resource the credential was theoretically allowed to reach.

This is an evidence problem first. A request count can tell you how busy a credential was. It cannot tell you whether the caller listed projects, read a secret, changed a webhook, or received a denial.

Counts mislead.

What did the leaked key actually touch?

Start by separating three quantities that incident notes often blur together:

  1. Entitlement radius: everything policy allowed the credential to do during the exposure window.
  2. Observed radius: resources and actions present in retained, attributable logs.
  3. Confirmed impact: successful operations whose semantics imply a read, mutation, deletion, or privilege change.

These sets answer different questions. Entitlements provide the conservative upper bound. Observed activity narrows the investigation. Confirmed impact drives recovery work, such as restoring a changed configuration or rotating a downstream secret that was read. A missing event is not proof that nothing happened; it may instead expose a retention, sampling, or instrumentation gap.

The exposure window needs equally careful bounds. Use the earliest credible disclosure or suspected access as the lower bound, add clock-skew tolerance appropriate to the systems involved, and stop at verified revocation. Preserve original timestamps and normalize a working copy to UTC. NIST's incident-handling guidance treats evidence collection, analysis, containment, and recovery as connected activities; that ordering is useful here because containment should not destroy the material needed for analysis.

Build a usage series before drawing the graph

The simple approach is to search for the literal key across logs. It fails in both directions: properly designed systems should not log raw secrets, while unrelated payloads can contain look-alike strings. OWASP recommends that secrets never be logged and describes rotation as replacing a secret without interrupting service. The useful correlation value is therefore a stable, nonreversible identifier generated at authentication time, not the credential itself.

A keyed digest is a practical identifier when the logging service holds a separate audit key. Record only a short display prefix if operators need one in dashboards, and retain the full digest in access-controlled evidence storage. Do not use an unkeyed hash for low-entropy credentials.

Each accepted or rejected request should produce a structured event with a UTC timestamp, credential fingerprint, principal or account, request ID, trace ID when available, source network context, normalized action, concrete resource identifier, decision, response class, and an audit-schema version. Mutation paths also need before/after references or a change-event ID. Authentication events alone are too shallow; database audit records or domain change logs provide the resource-level side of the join.

The focused example below consumes newline-delimited JSON that has already been exported into an evidence workspace. It selects one credential, applies explicit time bounds, sorts records, and summarizes distinct touches without pretending that a failed request changed anything.

from __future__ import annotations

import json
from collections import Counter, defaultdict
from datetime import datetime, timezone
from pathlib import Path


def instant(value: str) -> datetime:
    parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
    return parsed.astimezone(timezone.utc)


def reconstruct(path: Path, fingerprint: str, start: str, end: str) -> dict:
    lower, upper = instant(start), instant(end)
    events = []

    with path.open(encoding="utf-8") as stream:
        for line in stream:
            event = json.loads(line)
            occurred_at = instant(event["occurred_at"])
            if event.get("credential_fingerprint") != fingerprint:
                continue
            if lower <= occurred_at <= upper:
                events.append((occurred_at, event))

    events.sort(key=lambda item: (item[0], item[1]["event_id"]))
    resources = defaultdict(set)
    outcomes = Counter()

    for _, event in events:
        outcomes[event["outcome"]] += 1
        if event["outcome"] == "success":
            resources[event["action"]].add(event["resource_id"])

    return {
        "event_count": len(events),
        "successful_resources_by_action": {
            action: sorted(ids) for action, ids in sorted(resources.items())
        },
        "outcomes": dict(outcomes),
        "first_event": events[0][0].isoformat() if events else None,
        "last_event": events[-1][0].isoformat() if events else None,
    }
Enter fullscreen mode Exit fullscreen mode

This code is intentionally boring. It does not infer intent, and it does not label a timeout as a successful write. A production investigation should keep the ordered event rows beside the summary so every conclusion can be traced back to evidence.

Rotate without erasing the trail

Use a four-step state transition: issue, deploy, verify, revoke. Issue a replacement with the same or narrower policy. Deploy it through the normal secret-distribution path while both credentials are temporarily accepted. Verify that expected workloads authenticate with the new fingerprint and that the old fingerprint's traffic has drained. Then revoke the old credential and test that it is rejected.

No downtime is the goal, but overlap is still exposure. Keep it bounded and observable. The control-plane audit log should identify who created and revoked each credential, when the change occurred, and which account owned it. The data plane should never reveal either secret value.

There is a subtle trap here: a falling count for the old fingerprint does not prove every instance switched. Compare the new-key traffic against an expected workload inventory. Long-running workers, scheduled jobs, notebook kernels, and retry queues can remain quiet during the verification interval and reappear later. The deploy evidence needs to cover them, or their next authentication attempt must fail closed and alert clearly after revocation.

Preserve the leaked credential's fingerprint in the case record after revocation. Revocation changes authorization state; it should not make historical events unsearchable.

Correlate cautiously, then state confidence

Join gateway events to application and data records using request IDs or distributed trace context where available. W3C Trace Context Level 1 defines interoperable traceparent and tracestate headers, but a trace identifier is correlation metadata, not proof of identity. Authentication results remain the authority for which credential was presented.

For each resource, classify the evidence rather than forcing a binary verdict:

Finding Evidence threshold Response
Confirmed touch Attributable successful operation names the resource Recover or validate according to the operation
Attempted touch Attributable denial, validation failure, or unsuccessful operation Review policy and source; do not claim a mutation
Possible touch Gap, sampled telemetry, ambiguous correlation, or missing resource ID Expand the window and inspect adjacent systems
No observed touch Complete retained sources contain no matching event Report with retention and coverage limits

Keep inference out of the raw evidence layer. If a successful list operation returned ten records, the event schema should distinguish the collection that was queried from the item identifiers actually returned, when policy and privacy requirements permit that logging. If response bodies cannot be retained, document the weaker conclusion. Precision beats a dramatic incident summary.

This method has a hard limitation: it cannot reconstruct resource-level access that was never recorded. The trade-off is sharper privacy and lower audit-storage volume versus weaker incident certainty. Full payload capture is the wrong choice when prompts, source code, personal data, or secrets may appear in bodies; in that case, log typed resource identifiers and durable change references, then report any remaining ambiguity instead of widening collection by default. Teams without consistent resource IDs should fix that instrumentation boundary before treating the output as a definitive blast-radius report.

This is also where eval-driven development earns its keep. Treat reconstruction like an evaluation harness: seed synthetic audit fixtures for a read, a denied write, a duplicated delivery, an out-of-order event, and a missing trace link. The expected report should remain stable across parser changes. Keep fixtures free of real credentials and sensitive production payloads.

Measure this before copying the method

The method is only as good as its audit coverage. Before declaring it operational, measure the fraction of authenticated routes that emit a credential fingerprint and concrete action, the fraction of successful mutations linked to a durable change event, clock-offset bounds, duplicate-event behavior, and the searchable retention window. Also test how long it takes to answer one deliberately planted canary event from gateway entry to final resource. Use NIST SP 800-61 Revision 2 as incident-response background, while recognizing that an organization's legal, evidence-handling, and retention obligations may require a stricter process than this article describes.

For an AI feature, add token-bearing operations to the action taxonomy without logging prompts or model output by default. A useful event can record the authenticated account, model class, request ID, token counts returned by the serving layer, policy decision, and the IDs of retrieved documents. That supports cost and retrieval-scope analysis while avoiding a second leak inside the audit system. Any content capture needs a separate, explicit privacy and retention decision.

The decision rule is straightforward: use observed events for precise remediation, entitlements for the conservative search boundary, and telemetry gaps for uncertainty. Rotate through an auditable overlap, revoke decisively, and keep the evidence queryable. Anything more confident than the records allow is guesswork.

Sources

Top comments (0)