Short answer: treat the leaked API key as an identity under investigation, not as proof that every permitted resource was reached. In a B2B SaaS account platform, reconstruct its activity from four evidence layers: gateway access records, time-bucketed usage, application audit events, and downstream delivery outcomes. Preserve the raw records first. Then correlate by credential fingerprint, tenant, request or trace identifier, route template, and time window. The narrowest defensible blast radius is the intersection supported by those records; missing telemetry is an explicit unknown, never evidence of no access.
| Evidence layer | What it can establish | What it cannot establish alone | Pick this when |
|---|---|---|---|
| Gateway access records | A request associated with the key reached an ingress point | Whether the handler read or changed a particular object | You need request-level scope and timestamps |
| Usage time series | When activity rose, stopped, or diverged from a baseline | Which exact account, object, or field was touched | You need a fast incident window or a cross-check for missing logs |
| Application audit events | A business action was attempted or committed for a tenant or object | Every read, unless reads are deliberately audited | Auditability of account access is the primary decision axis |
| Delivery outcomes | An accepted platform event reached, retried, or failed at a downstream boundary | The caller's full behavior before or after that event | The backend had an outage and queued work may outlive the request |
Do not collapse these layers into one answer. A 200 at the gateway, a counter increment, an account-change record, and a delivered event describe different boundaries. The incident report should keep those boundaries visible.
What did the API key actually touch after the leak?
For access auditability, application audit events carry the most semantic weight because they can name the tenant, action, target type, target identifier, actor credential fingerprint, and outcome. They answer the question investigators actually have: what business data or control changed? Gateway logs are broader and often arrive sooner, so they define the search set. Usage series test the completeness of that set. Delivery records extend the timeline through an outage.
This creates a useful decision rule. Use gateway records to find candidate requests. Use audit events to prove object-level effects. Use downstream outcomes to follow asynchronous consequences. Use metrics to challenge the story when counts do not reconcile.
Counts can lie.
OWASP's Secrets Management Cheat Sheet recommends centralized auditing of secret access, including who requested a secret, what was requested, when, and whether access succeeded. It also warns that logs must not contain the secret itself. Apply both ideas to API credentials: retain a stable, nonreversible fingerprint for correlation, but never write the raw key into logs, spans, exception messages, or queue payloads. Rotation stops future use. It does not reconstruct past use.
Preservation comes before interpretation. Record the incident window, clock source, retention boundary, and export time. Restrict access to the evidence export and hash it so later analysis is reproducible. If retention expires during an investigation, a beautifully designed query still produces a false sense of precision.
Build a timeline that survives an outage
Suppose the account platform accepts customer events, validates the account, writes an inbox record, and dispatches work to the backend. During a backend outage, accepted events wait and retry. The original API request and the eventual account mutation can therefore occur in different minutes, perhaps on different hosts. Searching only the request window undercounts delayed effects. Searching only the recovery window loses the initiating identity.
The diagram in words is short: caller key -> edge request -> durable inbox -> worker attempt -> account audit event -> delivery receipt. Give the request an immutable correlation identifier at ingress and carry it through each durable handoff. Store a credential fingerprint beside the inbox record. A worker should copy those identifiers into its audit event even when it runs after the key has been revoked.
There are two clocks to retain: occurrence time from the originating component and ingestion time from the evidence store. Late delivery makes this distinction matter. Sort the human timeline by occurrence time, but use ingestion time to explain why an event was absent from an earlier snapshot. Keep timestamps in UTC and document the precision actually emitted by each layer.
Now reconcile counts across boundaries. If the gateway shows 1,042 accepted requests, the inbox shows 1,039 durable writes, and the audit stream shows 1,011 terminal outcomes, do not quietly label 31 requests harmless. The mismatch defines three investigation queues: requests without confirmed persistence, persisted items without a terminal outcome, and duplicate attempts that should map to one idempotent result. Those numbers are an example dataset, not a benchmark or an expected ratio.
A usage graph is valuable here, but only as a checksum with time attached. Aggregation destroys detail. A five-minute point containing 200 calls cannot tell you which 200 objects were read. It can reveal that an export missed records, that a retry storm extended beyond the suspected window, or that traffic continued after revocation.
Fast signal. Weak proof.
The minimum correlation contract
The implementation does not need a giant security schema. It needs a small contract used consistently at each boundary. The following TypeScript keeps raw credentials out of telemetry, templates routes to avoid identifiers in labels, and emits separate facts for ingress and business effects. The HMAC key used for fingerprinting is a protected investigation secret; separating it from API keys prevents a stolen API key from being used to calculate arbitrary fingerprints.
import { createHmac, randomUUID } from "node:crypto";
type Outcome = "accepted" | "denied" | "committed" | "failed";
type EvidenceEvent = {
occurredAt: string;
correlationId: string;
credentialFingerprint: string;
tenantId: string;
action: string;
targetType?: string;
targetId?: string;
outcome: Outcome;
source: "gateway" | "worker";
};
const fingerprint = (apiKey: string, auditKey: string): string =>
createHmac("sha256", auditKey).update(apiKey).digest("hex");
export function recordAcceptedRequest(input: {
apiKey: string;
auditKey: string;
tenantId: string;
routeTemplate: string;
}): EvidenceEvent {
return {
occurredAt: new Date().toISOString(),
correlationId: randomUUID(),
credentialFingerprint: fingerprint(input.apiKey, input.auditKey),
tenantId: input.tenantId,
action: `POST ${input.routeTemplate}`,
outcome: "accepted",
source: "gateway",
};
}
export function recordAccountEffect(input: {
accepted: EvidenceEvent;
accountId: string;
committed: boolean;
}): EvidenceEvent {
return {
occurredAt: new Date().toISOString(),
correlationId: input.accepted.correlationId,
credentialFingerprint: input.accepted.credentialFingerprint,
tenantId: input.accepted.tenantId,
action: "account.event.apply",
targetType: "account",
targetId: input.accountId,
outcome: input.committed ? "committed" : "failed",
source: "worker",
};
}
The sample returns records rather than choosing a logging library. That is deliberate. Send them to an append-oriented evidence store through the same telemetry path used by the application, with access controls and retention appropriate to incident investigation. Do not use targetId, tenant identifiers, credential fingerprints, or correlation identifiers as metric labels; their high cardinality belongs in event records. Metrics should count bounded dimensions such as route template, outcome, and source.
The schema also exposes a hard limit: accepted does not mean committed. If the durable write happens after the response, add another evidence event for persistence rather than renaming the gateway result. Vocabulary is part of the control. Ambiguous outcome names make incident reports overclaim.
Deployment deserves a staged check. Generate a synthetic key in a non-production tenant, send one accepted action and one denied action, then verify that the raw key appears nowhere while the same fingerprint and correlation identifier appear at the intended boundaries. Exercise a backend outage, restore the worker, and confirm that the delayed terminal event still joins to the ingress event. Finally, revoke the key and verify that a subsequent request is denied and recorded.
Turn records into a defensible blast radius
Start the investigation window before the earliest credible exposure and end it after revocation plus any queue drain. Preserve the original exports. Query all gateway events matching the leaked key's fingerprint, then join their correlation identifiers to inbox, worker, audit, and delivery records. Expand the window if usage counts disagree with event counts or if ingestion timestamps show late arrival.
Classify each candidate into one of four states:
- Confirmed effect: an audit record identifies the target and a committed outcome.
- Confirmed attempt: ingress or authorization records show the request, but no committed effect is supported.
- Pending or delayed consequence: durable work exists without a terminal record, including work held during the outage.
- Unknown: retention, sampling, clock error, or a missing correlation field prevents a defensible classification.
That fourth state matters. Zero matching rows can mean zero use, expired retention, the wrong fingerprint, an uninstrumented path, or an incomplete export. Call it unknown until independent evidence closes those alternatives. Absence of a record is useful only after telemetry coverage and retention have been established.
For reads, object-level proof may be harder than mutation proof. A database change log can support writes, but it will not generally prove every value returned to a caller. If read exposure matters, the application needs authorization and access audit events at the business boundary. Keep sensitive payload values out of those events. Record the target and decision, not the data itself.
The final incident artifact should separate confirmed targets, attempted targets, delayed work, and unknown coverage. It should also state how the key was revoked, which other credentials or derived tokens were rotated, and what evidence supports the end time. This is more honest than a single blast-radius number, and much easier for another reviewer to reproduce.
Limits worth stating plainly
Correlation cannot recover fields that were never recorded. Sampling weakens request-level claims. Shared API keys merge actors. Clock drift distorts ordering. A compromised application could also emit misleading telemetry, so high-risk investigations may require evidence from independently controlled boundaries. The trade-off is storage and operational complexity: request-level evidence costs more to retain and govern than a bounded metric series, while metrics cannot prove object-level access. This correlation approach is not appropriate as the sole evidence source when credentials are shared, audit events are sampled, or the application producing the records is inside the suspected compromise boundary. In those cases, use independently controlled gateway, identity, queue, and database records, and label any remaining gap unknown.
No telemetry creates certainty.
Keep the close concise: revoke and replace the credential, preserve evidence, trace durable consequences through recovery, and report claims at the resolution the records support. No stronger. Auditability is the ability to show the chain again to a skeptical reviewer.
Top comments (0)