DEV Community

YorkHolloway3257
YorkHolloway3257

Posted on

Leaked API Key Compromise: How to Rotate and Search Billing Logs

Short answer: treat a leaked API credential as both a security incident and an attribution incident. Record the report and the credential's immutable identity, revoke it, issue a replacement, then search append-only request records from the earliest plausible exposure through revocation. For an edtech platform that meters tutoring minutes or assessment runs per school, the invoice decision is strict: usage with a verified school identity may proceed; ambiguous usage goes into review rather than onto a customer's bill.

Rotation isn't closure. It stops future authentication with one secret, but it does not explain which school was charged, which requests succeeded, or whether a retry crossed the revocation boundary. The page should fire on suspected credential disclosure, unexpected use of a credential identity, or an impossible attribution state. A rising request graph by itself is weak evidence.

What should a leaked API key compromise report preserve?

Start the incident clock at the earliest plausible disclosure, not at the time someone opened the ticket. Preserve the original report, reporter, affected environment, detection time, possible exposure time, credential identifier, owning school account, and current responder. Never paste the secret into the ticket, chat, or log search. OWASP recommends that secrets have metadata such as an owner, creation time, expiration time, and rotation information, and that access and administrative actions be audited.

The credential identifier matters more than its display name. Names change; immutable IDs let responders join authentication events to metering records without handling secret material. Hashing a raw key for logs is still a poor default because low-entropy or structured credentials can become comparison targets. Authenticate at the edge, resolve the credential to a stable ID and school ID, and emit only those identifiers downstream.

For the billing case, capture 3 independent times: when the client says the lesson event occurred, when the edge accepted the request, and when the metering service committed it. The accepted time determines whether the old credential was still valid. The event time helps investigate delayed submissions. The commit time explains invoice inclusion. Collapsing them into one timestamp creates a postmortem argument instead of evidence.

One clock won't do.

Build an evidence record before the page fires

The safe implementation starts before an incident. Every accepted metered operation needs a durable event ID, credential ID, customer ID, quantity, authentication decision, request time, and disposition. Keep raw request logs access-controlled and retention-bound; keep the billing ledger append-only, with corrections represented as new records rather than overwritten history. OWASP's logging guidance warns against recording authentication passwords, access tokens, and encryption keys directly.

This Go example models the minimum evidence carried from authentication into billing. It uses a keyed HMAC to derive a searchable fingerprint from an already assigned credential ID; it does not hash or retain the secret itself. In production, the HMAC key belongs in a separately controlled secrets system and needs its own lifecycle.

package main

import (
    "crypto/hmac"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "time"
)

type MeterEvent struct {
    EventID       string    `json:"event_id"`
    SchoolID      string    `json:"school_id"`
    CredentialRef string    `json:"credential_ref"`
    Quantity      int       `json:"quantity"`
    Unit          string    `json:"unit"`
    AcceptedAt    time.Time `json:"accepted_at"`
    Decision      string    `json:"decision"`
}

func credentialRef(auditKey []byte, credentialID string) string {
    mac := hmac.New(sha256.New, auditKey)
    mac.Write([]byte(credentialID))
    return hex.EncodeToString(mac.Sum(nil))
}

func main() {
    e := MeterEvent{
        EventID:       "evt_01J9EDU7N4",
        SchoolID:      "school_1042",
        CredentialRef: credentialRef([]byte("replace-from-secret-store"), "cred_7f31"),
        Quantity:      45,
        Unit:          "tutoring_minute",
        AcceptedAt:    time.Date(2026, 10, 9, 2, 14, 7, 0, time.UTC),
        Decision:      "accepted",
    }
    b, err := json.Marshal(e)
    if err != nil {
        panic(err)
    }
    fmt.Println(string(b))
}
Enter fullscreen mode Exit fullscreen mode

Use a real random audit key, not the literal example value. Short-lived test fixtures are allowed to be obvious; production secrets are not.

Idempotency is part of attribution accuracy. If the client retries a 45-minute tutoring event during rotation, the same event ID must produce 1 billable ledger entry. A uniqueness constraint at the ledger boundary is stronger than a dashboard query that attempts to remove duplicates later. Reject records whose school identity cannot be resolved at acceptance time; do not infer the school from an IP address, email domain, or whichever account appears closest in time. This is a deliberate availability trade-off: an unattributed event may wait for review, while a guessed attribution can put another school's usage on the invoice and contaminate the very ledger needed to resolve the incident.

Contain first, but do not destroy the trail

Open the incident record and assign severity before changing the credential. Record the credential ID and revocation start time, revoke the exposed credential, issue a distinct replacement, and distribute it through the normal secret-delivery path. Do not edit the old database row into a new key. Separate identities preserve the boundary between old and new activity.

There is a trade-off. Immediate revocation can interrupt classrooms; overlap can extend an attacker's window. For confirmed public disclosure, revoke first. For suspicion without evidence of disclosure, a tightly bounded overlap may be justified only when the service owner explicitly accepts it, both credential IDs remain distinguishable, and an automatic expiry is set. The responder should write that choice into the incident timeline. Silence is not approval. OWASP recommends automating rotation where possible and designing consumers to handle rotation without extended outages. Automation still needs a failure state the pager can understand: old credential revoked, replacement created, consumer updated, and a synthetic authenticated request accepted under the replacement. A single green “rotation complete” badge can hide which step failed, so the incident record must retain the result of each transition rather than only the final status.

Revoke means revoke.

Search for billing blast radius, not merely traffic

Search from the earlier of the suspected disclosure time and the first anomalous use, ending only after revocation has propagated through every verifier. Partition results by credential ID, school ID, decision, route class, event ID, and time bucket. Then reconcile accepted metering events against the billing ledger.

The useful questions are concrete: Did the old credential authenticate after the recorded revocation time? Did 1 event ID appear with 2 quantities? Did a credential owned by school_1042 create usage for another school? Did rejected requests enter the ledger? Which accepted events lack a corresponding immutable billing record?

The following standalone Go program reads newline-delimited JSON exported from the controlled audit store and flags attribution conditions. It deliberately avoids printing secrets or full request bodies.

package main

import (
    "bufio"
    "encoding/json"
    "fmt"
    "os"
    "time"
)

type Record struct {
    EventID        string    `json:"event_id"`
    SchoolID       string    `json:"school_id"`
    CredentialID   string    `json:"credential_id"`
    CredentialOwner string   `json:"credential_owner"`
    Accepted       bool      `json:"accepted"`
    Ledgered       bool      `json:"ledgered"`
    Quantity       int       `json:"quantity"`
    AcceptedAt     time.Time `json:"accepted_at"`
}

func main() {
    revokedAt, err := time.Parse(time.RFC3339, "2026-10-09T02:20:00Z")
    if err != nil {
        panic(err)
    }

    seen := map[string]int{}
    s := bufio.NewScanner(os.Stdin)
    for s.Scan() {
        var r Record
        if err := json.Unmarshal(s.Bytes(), &r); err != nil {
            fmt.Fprintf(os.Stderr, "invalid record: %v\n", err)
            continue
        }
        if r.CredentialID == "cred_7f31" && r.Accepted && r.AcceptedAt.After(revokedAt) {
            fmt.Printf("accepted_after_revocation event=%s\n", r.EventID)
        }
        if r.Accepted && r.SchoolID != r.CredentialOwner {
            fmt.Printf("owner_mismatch event=%s\n", r.EventID)
        }
        if !r.Accepted && r.Ledgered {
            fmt.Printf("rejected_but_ledgered event=%s\n", r.EventID)
        }
        if old, ok := seen[r.EventID]; ok && old != r.Quantity {
            fmt.Printf("quantity_conflict event=%s\n", r.EventID)
        }
        seen[r.EventID] = r.Quantity
    }
    if err := s.Err(); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

Don't convert every flagged row into fraud. Clock skew, queued work, and partial deployment are competing explanations, which is why the 3 timestamps and verifier identity belong in the evidence. The program is a triage tool; the ledger and authentication records remain authoritative.

Bill only the unambiguous set. Quarantine events with owner mismatches, quantity conflicts, acceptance after confirmed revocation, or missing ledger linkage. That may delay part of an invoice, but it avoids making a school finance team finance your uncertainty. The incident owner can later release, credit, or void those events through a documented adjustment.

Verify containment and prepare rollback

Containment is verified when the old identity is rejected at every verifier, the new identity passes a synthetic request, duplicate event IDs remain single ledger entries, and the review queue contains every ambiguous event in the selected window. Query the underlying records. Dashboards aggregate, sample, and relabel; they are navigation aids, not evidence that an invoice is correct.

Ask what page fired. A useful page identifies an actionable failure such as “revoked credential accepted” or “meter event owner mismatch,” links to the runbook, and includes credential and customer references without secret material. A page that says only “API traffic unusual” leaves the responder to rediscover the incident model at 3 a.m.

Rollback does not mean reactivating the leaked credential. If the replacement breaks consumers, roll back the consumer configuration or deployment while keeping the compromised identity revoked; if service continuity requires another credential, issue another distinct one. Preserve failed rotation attempts in the audit trail. Afterward, test the runbook with a synthetic credential, confirm retention covers the billing dispute window required by policy, and add postmortem actions for every missing field or manual join.

Close the incident only after security containment and billing reconciliation have separate owners and explicit completion criteria. One can finish before the other. The durable outcome is not a rotated string; it is a defensible account of which school usage was accepted, rejected, held, and ultimately invoiced.

References

Top comments (1)

Collapse
 
alewx profile image
alex •

how do you keep the log scan performant when you have millions of immutable rows and retain them for months, and does revocation propagate instantly across all services or is there a window where the leaked key still works? if you're depending on eventual consistency you'll need a fallback to reject stale usage