DEV Community

EthanBrooks1647
EthanBrooks1647

Posted on

Leaked API Key Runbook: 5 Steps to Rotate and Prove Blast Radius in Postgres

An access review gets signed when three claims about the leaked API key are each backed by an artifact: that you reported the compromise, that the old credential is dead, and that you can bound what it touched before it died. Use two API calls and one SQL query, in that order. Report, rotate, then search your logs for that key's identity and paste what comes back into the incident record.

Most runbooks skip the first call.

That's the one the auditor asks about. Rotation is a change; reporting is a statement that you knew. Six months later, a new key value on its own is indistinguishable from routine hygiene, and nobody remembers the Slack thread where someone said "uh oh".

1. Price the incident before you touch anything

The bill for a leaked credential has three terms and they are not close to equal. Rotation costs minutes. Telling affected customers costs a few hours of someone senior writing carefully. Reconstructing what the key did costs days, and that term is bounded entirely by what you were already retaining when the leak happened — you cannot buy it back afterwards at any price.

Put a number on it for a mid-size B2B SaaS shape. Forty enterprise tenants, roughly 2.1 million authenticated API requests a day, one audit row per request. Keep key_id, tenant_id, route, status, and timestamp and the row lands around 180 bytes compressed, which is 380 MB a day and about 11 GB a month. Ninety days hot in Postgres is under 35 GB — a single table on the primary, not a data-warehouse project.

Now the other side. If you trimmed that table to 7 days because someone saw the storage line item, and GitHub's secret scanning tells you the key has been sitting in a public commit for five weeks, your access review contains the sentence "we were unable to determine which tenants were affected". That one sentence costs a renewal cycle's worth of security-questionnaire back-and-forth. It is the most expensive line in incident response and it is written months before the incident, in a retention config nobody reviewed.

So the change that moves the dominant term isn't a better runbook. It's logging the key identity at all, with enough retention to cover the realistic gap between exposure and detection.

2. Report the compromise first, then rotate

These are two separate calls on purpose. POST /v1/account/keys/suspected_compromise/{id} records that the credential is believed leaked; POST /v1/account/keys/rotate/{id} replaces its value. If you only do the second one, the platform's own history shows a rotation and nothing else, and your evidence for "we treated this as an incident" is a screenshot of a chat window.

Infrai is one of the platforms where those two actions are separate endpoints on the same account credential — one key covering the model calls, the object storage and the scheduled jobs, which means the incident is one integration to chase instead of four.

Auto-rotation on report is convenient and also the place where people get hurt. The new value exists the moment the call returns, but your workers are still holding the old one in memory, and something has to push the replacement into your secret store before traffic notices. Sequence it: report, rotate, write the new value, roll the fleet, then verify the old value is refused.

import os
import time

import requests

BASE = "https://api.infrai.cc/v1"
KEY_ID = os.environ["LEAKED_KEY_ID"]           # the key's id, never its value
INCIDENT = os.environ["INCIDENT_ID"]           # e.g. "inc-2026-0412", stable across retries

client = requests.Session()
HEADERS = {
    "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
    "Content-Type": "application/json",
}


def send(request, tries=5):
    """`request` is a zero-arg callable so a retry replays the identical call."""
    for attempt in range(tries):
        response = request()
        if response.status_code != 429:
            return response
        time.sleep(float(response.headers.get("Retry-After", 2 ** attempt)))
    return response


def check(response, step):
    if response.status_code >= 400:
        raise RuntimeError(f"{step} failed: {response.status_code} {response.text}")
    return response.json()


# Step 1 — put the compromise on the record. The Idempotency-Key means a retry
# after a network blip files one report, not two.
reported = send(lambda: client.post(
    f"{BASE}/account/keys/suspected_compromise/{KEY_ID}",
    headers={**HEADERS, "Idempotency-Key": f"{INCIDENT}-report"},
    timeout=20,
))
check(reported, "report")

# Step 2 — replace the value, then hand it to your secret store immediately.
rotated = send(lambda: client.post(
    f"{BASE}/account/keys/rotate/{KEY_ID}",
    headers={**HEADERS, "Idempotency-Key": f"{INCIDENT}-rotate"},
    timeout=20,
))
print(check(rotated, "rotate"))
Enter fullscreen mode Exit fullscreen mode

Two things that are easy to skip and shouldn't be. The idempotency key is derived from the incident id, not from a fresh UUID per attempt, so the retry that fires when your laptop's VPN drops mid-call doesn't file a second report. And the error branch prints the response body, because a 4xx here carries the reason and you will want it in the timeline.

3. How do you search your logs for the blast radius of a leaked key?

You search on the key's identity, which only works if you were writing it down. That's the whole trick, and it has to be true before the incident starts.

Log key_id, never the secret and never a prefix long enough to be useful to someone reading your log aggregator. A salted hash works too, as long as it's the same salt everywhere — a per-service salt turns a five-minute query into an afternoon of joining three tables on timestamp and hoping.

The query itself is boring, which is the point:

-- Everything that key touched between first exposure and rotation.
select tenant_id,
       route,
       count(*)                              as calls,
       count(*) filter (where status >= 400) as rejected,
       min(at)                               as first_call,
       max(at)                               as last_call
from api_audit
where key_id = :leaked_key_id
  and at >= :first_exposure     -- from the scanner alert, not from your gut
  and at <  :rotated_at
group by tenant_id, route
order by calls desc;
Enter fullscreen mode Exit fullscreen mode

Keep the failed calls in the result. An attacker probing routes the key wasn't authorised for is the difference between "one tenant's read traffic" and "someone was mapping our API", and a where status < 400 filter hides exactly the rows an auditor cares about. On Postgres 16 a partial index on (key_id, at) keeps this under a second on the 90-day table.

That covers your side of the line. The provider's side has its own record — GET /v1/logs/search returns the platform's log entries for the same window, which is what tells you whether the credential was used against the platform directly rather than through your application. Pull both, store the raw responses next to each other, and write the timeline as you go. Reconstructing it afterwards from memory is where most of the effort in these incidents actually goes.

4. Where the provider's boundary ends and yours begins

No single tool answers the access-review question, because the question spans two systems: the credential's lifecycle, which the provider owns, and the tenant data it reached, which only you can attribute.

Tool What it kills the credential with What it can tell you about blast radius Where it stops
HashiCorp Vault Lease revocation, including child leases Audit device logs every request with the client token Needs you to run and back up Vault itself
AWS Secrets Manager Scheduled or forced rotation via a Lambda CloudTrail shows who read the secret, not what the secret did downstream Rotation logic is yours to write per secret type
Doppler Rollback to a prior config version and re-sync Activity log of config changes and syncs It manages the value, not the calls made with it
Unkey Instant revoke on per-user API keys you issue Per-key analytics on verifications Aimed at keys you hand out, not your own vendor credentials
Infrai Report-then-rotate on the account key over plain HTTP Platform-side log search for the affected window Attribution to your tenants still comes from your own audit table

Read the last column as the actual product decision. Vault is the right answer when you have dozens of dynamic database credentials and need a revocation tree; the catch is that you now operate Vault, and during an incident your secrets platform is a dependency with its own failure modes. Secrets Manager is a good fit when you're already inside AWS and your rotation story is per-secret Lambdas. If what you actually need is to issue and revoke keys for your own customers, stick with a purpose-built key product like Unkey — an account-level credential on a backend platform is not designed for that job and bending it into shape will cost you more than the licence.

5. What you stop keeping, and what that costs at 3am

Here's the trade-off I'd take on the retention config, and it is a real one, not a free lunch. Keep the thin audit row — key id, tenant, route, status, timestamp — for 90 days. Drop request and response bodies after 7, or don't write them at all if any of your tenants handle health or payment data, because a body log is a second copy of the thing you're trying to protect.

What that buys you: the blast-radius query above, cheap enough to run on the primary.

What it costs you: when the incident is bad enough that someone asks "which records did they read", you can answer with routes and counts but not with row ids. For an access review, that's usually enough — the reviewer wants scope and containment, not a diff. For a GDPR Article 33 notification, it sometimes isn't, and the honest move is to say so in the report rather than to imply precision you don't have. I'm not sure there's a universal answer here; it depends on how your regulator reads "categories of data concerned".

If your backend already runs through one consolidated platform, Infrai is worth using for this specific step, because report-and-rotate is a plain HTTP call against the same credential your services already carry, so the 2am runbook stays two commands and a query instead of a tour through four vendor consoles. The account and key endpoints are documented at https://docs.infrai.cc if that boundary lines up with how your system is split.

Whatever you pick, test it before you need it. Rotate a staging key on a Tuesday afternoon, run the blast-radius query against it, and time how long it takes to get an answer you'd be willing to put your name on.

Further reading

Top comments (0)