TL;DR: For a leaked property-management API credential, choose an append-only access-event stream joined to a per-credential usage series. The stream proves which tenant, property, and operation were reached; the series reveals the time window and gaps worth investigating. A usage graph alone is cheaper to retain and faster to scan, but it cannot establish blast radius.
| Choice | Identifies touched resources | Shows timing and volume | Audit value | Burden |
|---|---|---|---|---|
| Correlated events plus usage series | Yes | Yes | Highest | Higher |
| Access events only | Yes | Partly | High | Medium |
| Usage series only | No | Yes | Low | Lowest |
Recommendation: retain both views, generated from the same credential identity and correlation ID. During the drill, treat resource-level access events as evidence and the usage series as the index that tells you where to look. This costs more storage and implementation time. Auditability is the point.
What can logs establish after an API key leak?
Start with the question an operator must answer after rotation: during the suspected interval, did this credential read tenant contact data, change a lease, retrieve a maintenance attachment, or only hit a health check? Request totals cannot distinguish those outcomes. A resource access event can, provided it records stable identifiers and the authorization result.
The minimum useful event includes event time, credential ID, actor or service identity, tenant ID, property ID when applicable, action, resource type, a pseudonymous resource ID, authorization outcome, response class, source-network context, and a correlation ID. Never put the secret itself in the event. OWASP's secrets guidance calls for auditing who requested a secret, what secret was requested, and whether access was granted, while warning that logs must not contain the secret value.
Keep the credential ID stable across rotation records, but issue a new version identifier for the replacement. Otherwise a graph can silently merge pre-rotation and post-rotation traffic. Record denied attempts too. A denied export did not expose tenant data, yet it matters when reconstructing intent.
This is the first criterion: preserve authorization context at resource granularity without preserving credentials or sensitive payloads. Store identifiers that support a join, not copies of leases, phone numbers, or maintenance photos.
Correlation is the second criterion
A drill falls apart when three clocks and four identifiers disagree. Normalize timestamps to UTC, retain the original event time, and separately record ingestion time. The difference makes late delivery visible. Pass a request correlation ID through the gateway, authorization decision, application event, and usage aggregation.
Short gaps matter.
Count them.
Suppose the usage series reports 41 requests in a five-minute bucket while the event query returns 39 terminal outcomes. That is not proof that two requests touched data. It is a reconciliation failure with several possible states: queued, dropped, duplicated, or outside the query boundary. Never round that mismatch into a confident incident statement.
For a small service, I would define one invariant before adding dashboards: every counted request must map to exactly one terminal access outcome, including rejection and internal failure. A daily check can compare counts by credential version and time bucket. Outsource the undifferentiated storage and charting if needed, but keep the event schema, retention policy, and reconciliation test under your control. Those determine what you can prove.
A drillable TypeScript evidence path
This interface emits a sanitized event and derives usage from the same record. The sink may be a file, queue, database, or managed log service; the contract does not depend on one.
type Outcome = "allowed" | "denied" | "failed";
type AccessEvent = {
occurredAt: string;
ingestedAt: string;
credentialId: string;
credentialVersion: string;
correlationId: string;
tenantId: string;
propertyId?: string;
action: "read" | "create" | "update" | "delete" | "export";
resourceType: "lease" | "resident" | "work_order" | "attachment";
resourceRef: string;
outcome: Outcome;
};
interface EvidenceSink {
append(event: Readonly<AccessEvent>): Promise<void>;
incrementUsage(event: Readonly<AccessEvent>): Promise<void>;
}
async function recordAccess(
sink: EvidenceSink,
event: Omit<AccessEvent, "ingestedAt">,
): Promise<void> {
const complete = { ...event, ingestedAt: new Date().toISOString() };
await sink.append(complete);
await sink.incrementUsage(complete);
}
The two awaited calls are not atomic. Use a durable transaction or an outbox with idempotent consumers, then verify the invariant during the drill. The example demonstrates a shared source record, not a complete delivery guarantee.
That limitation is real.
Seed the exercise with synthetic identifiers and a known timeline: allowed reads across two properties, a denied export against a third, and a late-arriving work-order update. Rotate the credential, retire the old version, then query the evidence without consulting the seed sheet. Separate confirmed access, confirmed denial, ambiguous gaps, and activity after rotation.
A compact finding could read: "Credential version v7 made 18 allowed requests against properties p-14 and p-22 between 02:10 and 02:17 UTC; one export against p-31 was denied; zero old-version events were observed after rotation; two counted requests lack terminal events and remain unresolved." Every clause maps to evidence. No guesswork.
Run the drill like an incident
Freeze the investigation window before browsing events. Record the first suspicious observation, earliest available event, rotation time, and last old-version attempt. Export an immutable investigation set and hash it. Restrict access because metadata can still reveal tenant relationships even when payloads are absent.
Then follow a short sequence:
- Disable and rotate the exposed credential while preserving its non-secret identifier.
- Capture the time bounds and evidence-retention bounds.
- Reconcile usage buckets with terminal access events.
- Enumerate allowed operations by tenant, property, resource type, and action.
- Separate denied, failed, duplicate, late, and unresolved records.
- Test the retired version and confirm the attempt is rejected and recorded.
- Save the query, evidence hash, assumptions, and unresolved gaps.
The useful success condition is reproducibility: a second operator, or future you, can run the saved query against the frozen set and reach the same boundary. Ship weekly; rehearse the path that can interrupt a week.
When is the runner-up enough?
Access events without a separate usage series can be better for a low-volume internal integration. If event queries return quickly, retention is adequate, and reconciliation proves continuity, deriving charts on demand avoids another pipeline. The loss is fast anomaly scanning, not resource attribution.
The trade-off is operational: correlated events demand more storage, stricter schema ownership, and reconciliation work. This approach is not suitable when the system cannot retain resource-level metadata under its privacy policy; use coarser tenant-level events instead, document the weaker boundary, and do not claim resource attribution.
Usage series alone is acceptable only as a narrow signal that a supposedly idle credential became active. It can trigger an investigation. It cannot close one, because ten requests might mean ten health checks, one resident export retried nine times, or access across ten properties.
The decision rule is blunt: if the incident report must name affected tenants or resources, aggregated usage cannot be the system of record. Choose the smallest architecture that preserves that claim, test its gaps, and set retention according to how long a leak could plausibly remain undiscovered.
Sources
- OWASP, "Secrets Management Cheat Sheet": https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
Top comments (0)