DEV Community

RaffertyBarrett4726
RaffertyBarrett4726

Posted on

Compromised API Key Reporting vs Quietly Rotating: Preserve the Traffic Record

Report a suspected compromised API key through the incident channel, preserve the relevant evidence, and then revoke it promptly. Quiet rotation is appropriate for routine maintenance, not for a leaked-key drill. Short answer: choose the reported path when exposure is plausible; the incident record lets the team distinguish unauthorized use from the legitimate requests it may refuse during containment. Reporting is not permission to leave a known exposed credential active while paperwork catches up.

In a B2B SaaS service, one integration key may authenticate a customer's scheduled exports and queue workers. Revoking it can stop the leak and refuse legitimate traffic at the same time. The spend ceiling matters, but it cannot answer whether an unexpected burst was an attacker, a retry storm, or a customer running a backfill. That distinction needs a record made before context disappears.

What does reporting a compromised API key preserve that quietly rotating loses?

A replacement credential restores access for a cooperating client. It does not establish when the old credential became exposed, which principal used it, whether it reached unexpected resources, or how many requests were rejected after revocation. OWASP's Secrets Management Cheat Sheet treats rotation, revocation, monitoring, and auditability as related parts of managing secrets. An incident record makes those observations available to responders who were not on the first call. If the export scheduler crosses a revocation boundary, the owner needs to reconcile scheduled work against completed work and refused requests, not just inspect an authentication error graph; each of those three counts answers a different question.

Keep the timeline.

The operational trap is a queue. If a scheduled export fails authentication, its retry may look like fresh malicious traffic unless the team correlates the attempt with a job identifier and the time the credential was revoked. Conversely, a successful request before revocation is not proof that it was authorized merely because it used a valid key. Record the distinction; do not guess from the status code.

Quiet rotation can also erase a useful ordering of events. Capture the discovery time, reporter, affected credential identifier, observed usage window, last known legitimate use, revocation time, replacement deployment time, and decision owner. Store a fingerprint or identifier, never the full secret in the ticket. Keep timestamps in UTC and distinguish observed times from estimates. The record should say what remains unknown.

Set the ceiling before the drill starts

Treat the spend ceiling and refused traffic as separate constraints. A ceiling is a containment control for additional consumption, not a substitute for revocation; an overly broad block can interrupt every tenant sharing an integration path. Before exercising the drill, write down which credential can be revoked, which tenant and workloads use it, who can approve the action, and how the team will detect refused legitimate requests. If attribution is uncertain, bound the exposure with the narrowest available control while the incident owner checks the evidence, then revoke the suspect key without waiting for perfect certainty.

For a rehearsal, use explicit hypothetical thresholds, not claimed production measurements: an alert at 80% of an approved test budget, and a decision checkpoint after 15 minutes of unexplained use. These numbers are drill inputs. They are not universal safe limits. Set the real threshold from your own workload, response time, and tolerance for refused exports; write down who may change it during an incident.

The useful trade-off is visible in a short decision log: "revoked at 14:05 UTC; 12 scheduled jobs still held the old key; reissued to the integration owner; queue retries paused until credentials were updated." Those are illustrative entries, not an incident report. The log should also capture the alternative considered and why it was rejected. This is the difference between a reversible containment decision and a quiet credential change that leaves the next shift to reconstruct intent. The easier path may restore the first worker quickly, but it provides no shared explanation for the other 11 jobs when their retries begin. The refusal count needs a time window and a tenant boundary; an aggregate error count can hide both a stalled customer and unrelated healthy traffic.

Run containment without multiplying deliveries

Start by assigning an incident owner and preserving access logs, audit events, relevant job IDs, and the current policy configuration under your retention rules. Restrict access to that evidence. Notify the credential owner through the established incident channel, mark the credential compromised, and revoke or disable it. Issue a replacement through the normal secrets workflow, update dependent workloads, and verify the old credential no longer authenticates. Never paste it into a chat message or test request.

Pause affected queue consumers if retries would create duplicate business actions or mask the authentication failures. On resume, each export should carry a stable operation ID so a retry can be recognized as the same operation; a fresh credential does not make the work itself idempotent. Count both refused requests and completed exports by tenant and operation ID. A fall in authentication errors alone may mean workers stopped trying, not that customers recovered.

This compact Go record captures the decision boundary without putting a secret in the incident log. The caller must supply an opaque credential ID, not the key itself; the record should be written to the team's access-controlled incident store.

package incident

import "time"

type KeyDecision struct {
    CredentialID string
    TenantID     string
    OperationID  string
    ObservedAt   time.Time
    RevokedAt    time.Time
    Decision     string
    Reason       string
}

func RecordRevocation(credentialID, tenantID, operationID, reason string, observed, revoked time.Time) KeyDecision {
    return KeyDecision{
        CredentialID: credentialID,
        TenantID:     tenantID,
        OperationID:  operationID,
        ObservedAt:   observed.UTC(),
        RevokedAt:    revoked.UTC(),
        Decision:     "revoke",
        Reason:       reason,
    }
}
Enter fullscreen mode Exit fullscreen mode

Don't log the key.

Separate investigation from restoration. Inspect the preserved pre-revocation window for unexpected source, tenant, operation, and volume patterns. Do not infer intent from IP address alone. If evidence suggests broader access, escalate according to your incident process and preserve the original observations; rotating again without documenting the reason only fragments the timeline. OWASP's incident response guidance calls for a defined process and evidence handling, which is why the report stays open after the credential is replaced.

Verify recovery and define rollback

Exercise a known authorized export with the replacement key, confirm that the old key is refused, and check that queued work completes once per operation. Compare the count of scheduled jobs, attempted deliveries, refused authentication attempts, and successful outputs over the same time window. Record gaps for investigation. Also check that the ceiling still caps additional consumption as intended without blocking unrelated tenants.

Do not roll back by reactivating the exposed credential. If the replacement rollout breaks legitimate traffic, roll back the deployment or configuration that distributed the replacement, pause retries, and issue another fresh credential through the approved path. Keep the compromised key disabled. Close the drill only after the owner signs off on evidence retention, customer impact, follow-up notifications where required, and the unresolved questions. The measure of success is not merely a new secret: it is a defensible sequence of containment and restored service.

No silent reopen.

References

Top comments (0)