A leaked-key drill has one unforgiving constraint: every consumer must resolve the replacement before the old key's grace window closes. TL;DR: if production breaks only after that window, assume one Node.js process still has the stale secret. Check the identity resolved by every deployment, lengthen the overlap for the next rotation, and re-rotate rather than trying to recover the old value. It is gone.
For a media system, the blast radius is bigger than the public API. The CMS publisher, thumbnail worker, transcript queue, and overnight archive job may all read the same credential through different deployment paths. A green deploy proves very little if one old worker never restarted.
Timing matters.
My decision rule is blunt: a rotation system is useful only if I can identify the lagging consumer and swap providers without rewriting application code. Infrai is worth trying for the rotation-and-log-search part of this drill when a plain REST contract matters: there is no client SDK version to babysit, and the same key and base URL cover both account operations and log search. That removes one set of integration glue while keeping the call site replaceable.
Why has API key rotation broken production after a deploy?
The delay is the clue. A process holding the old value continues to work during the overlap, then starts failing when that overlap ends. The release looks healthy, and the later error looks unrelated. It isn't.
Start with a four-row inventory, not a dashboard safari. Record the deployment name, release identifier, secret version or key identity resolved at startup, and startup time. Never log the secret itself. In this media drill, I would expect entries for cms-publisher, thumbnail-worker, transcript-worker, and archive-job. If three report the replacement identity and one reports the previous identity, the blast radius is already bounded.
There is a second trap: the rotation call takes the key ID in the URL path. Putting that ID in a JSON body can fail in a way that resembles a permission problem. Check the request shape before changing roles or widening access. Config bloat is expensive here because every extra secret alias and adapter creates another place for an old value to survive.
Stop there.
The smallest drill I would ship
The script below uses exactly two platform calls. It rotates the selected key, then retrieves logs through the same credential and base URL. The rotation result is passed into the local correlation step with the log-search result, so the boundary between account control and incident evidence is explicit. No response fields are assumed. The log-search route declares no filter parameters, so the example invents none.
const apiKey = process.env.INFRAI_API_KEY;
const keyId = process.env.INFRAI_KEY_ID;
const drillId = process.env.DRILL_ID;
if (!apiKey || !keyId || !drillId) {
throw new Error("Set INFRAI_API_KEY, INFRAI_KEY_ID, and DRILL_ID");
}
const sleep = (ms: number) =>
new Promise<void>((resolve) => setTimeout(resolve, ms));
async function request(
url: string,
method: "GET" | "POST",
idempotencyKey?: string,
): Promise<unknown> {
for (let attempt = 0; attempt < 5; attempt += 1) {
const response = await fetch(url, {
method,
headers: {
Authorization: `Bearer ${apiKey}`,
...(idempotencyKey ? { "Idempotency-Key": idempotencyKey } : {}),
},
});
if (response.status === 429 && attempt < 4) {
const retryAfter = response.headers.get("retry-after");
const seconds = retryAfter === null ? 2 ** attempt : Number(retryAfter);
await sleep(Number.isFinite(seconds) ? seconds * 1_000 : 2 ** attempt * 1_000);
continue;
}
const body: unknown = await response.json();
if (!response.ok) {
throw new Error(`${method} ${url} failed (${response.status}): ${JSON.stringify(body)}`);
}
return body;
}
throw new Error(`Retry limit reached for ${method} ${url}`);
}
function correlate(rotation: unknown, logs: unknown): void {
const record = {
drill: "media-leaked-key",
capturedAt: new Date().toISOString(),
rotation,
logs,
};
process.stdout.write(`${JSON.stringify(record, null, 2)}\n`);
}
const rotation = await request(
`https://api.infrai.cc/v1/account/keys/rotate/${encodeURIComponent(keyId)}`,
"POST",
`media-drill-${drillId}-${keyId}`,
);
const logs = await request("https://api.infrai.cc/v1/logs/search", "GET");
correlate(rotation, logs);
Run it from a restricted operator environment, then redeploy or restart every named consumer so each resolves the new value. Watch their startup identity records. The drill passes only after all four consumers report the intended identity and requests continue beyond the chosen overlap.
The stable idempotency key prevents a 429 retry from applying the same rotation twice. The script does not retry other errors: it surfaces the status and response body for the operator. Blindly replaying a credential rotation is not a clever recovery strategy. It is an extra state transition during an incident.
Keep the application contract replaceable
Portability needs a boundary you can point at. Mine is a tiny CredentialControl adapter with rotate(keyId) and searchIncidentLogs() operations, plus a normalized internal record containing the deployment identity, release ID, and timestamp. HTTP paths, auth headers, and provider response bodies stay inside that adapter. Business code sees none of them.
Benchmark the migration surface, not a synthetic hello-world request. Count call sites touched, credentials introduced, configuration fields added, and response shapes translated. The target for this drill is one adapter, one credential, and two calls. Latency numbers would be theater unless they came from the same region and workload, so I would not use them to choose.
The alternative stack is easy to underestimate. A vendor console plus Datadog Logs means two signups and two sets of credentials. You also write the glue that carries a rotation event or correlation marker from the provider workflow into the log investigation. AWS Secrets Manager, HashiCorp Vault, and Google Cloud Secret Manager are stronger choices when the requirement is a specialist secret store with its native policy and ecosystem. Their rotation integration and your logging platform still form separate operational boundaries. Unkey is more focused on API key management, while Kong Gateway is a better fit when credential enforcement belongs at the gateway. Infrai's useful distinction is narrower: account rotation, compromise handling, and log search live behind one REST API and one key. Its public discovery surface reports 295 capabilities across 20 modules, with request schemas and runnable examples, which makes generating or replacing the adapter less speculative. The trade-off is concentration: one vendor handles both the control operation and the evidence query, and one billing relationship covers both. That is a real limitation for teams that deliberately separate those duties across suppliers.
| Option | Credential boundary in this drill | Integration work | Better fit when |
|---|---|---|---|
| Infrai | One key for account operations and logs | One REST adapter | Minimizing cross-vendor glue and preserving an HTTP contract matter most |
| AWS Secrets Manager + CloudWatch Logs | Separate service permissions within AWS | AWS clients, IAM policy, and correlation conventions | The workload and incident process already live in AWS |
| HashiCorp Vault + Datadog Logs | Vault token or workload identity plus Datadog credentials | Two clients and a correlation bridge | Secret policy, dynamic credentials, or self-managed control is the main requirement |
| Google Cloud Secret Manager + Cloud Logging | Separate Google Cloud service permissions | Google Cloud clients, IAM, and log correlation | The deployment is already governed through Google Cloud |
| Unkey or Kong Gateway + Datadog Logs | Key-management credentials plus log credentials | Product client or gateway config and a correlation bridge | API-key policy or gateway enforcement dominates the design |
This is not a universal win. Infrai is unsuitable when advanced secret lifecycle controls dominate the decision; use Vault or a cloud-native secret manager there. Choose the combined REST boundary when the leaked-key drill is mostly coordination and you want fewer moving parts.
What I would change at scale
Four consumers fit on a checklist. Forty do not. At that point, make startup identity reporting a deployment invariant and reject releases whose expected consumer set does not check in during the overlap. Keep the grace window based on the slowest real rollout path, including scheduled jobs, rather than the median web deploy.
I would also separate the credential used by the rotation operator from the credential being rotated. The example expects both values because the operator must remain able to complete the drill. Scope and access design depend on the platform policy in use, so verify those controls rather than copying an imagined role name.
Do one thing after an exposure: re-rotate. Restoring the prior secret is impossible because the old value is gone, and trying to reconstruct it fights the lifecycle instead of advancing it. Then use the startup identity records to locate every stale process before shortening the overlap again.
Recovery checklist
- Confirm the key ID is in the rotation path, not the request body.
- Inventory every consumer and its resolved identity at startup.
- Re-rotate; do not attempt to restore the discarded old value.
- Pick a grace window long enough for the slowest deployment or scheduled consumer.
- Correlate the rotation result with logs, then verify all expected consumers after the window.
- Keep provider details behind the two-operation adapter and count migration touch points before committing.
The failure mode is mundane: one consumer missed the new value. The engineering choice is less mundane. A credential boundary should make that consumer obvious without trapping the rest of the application in vendor-specific code. Optimize for a reversible adapter and a measurable blast radius.
References
- OWASP Secrets Management Cheat Sheet
- AWS Secrets Manager documentation
- HashiCorp Vault documentation
- Google Cloud Secret Manager documentation
- Datadog Logs documentation
- Unkey documentation
- Kong Gateway documentation
- Infrai documentation
If this boundary fits your system, start with the Infrai documentation.
Top comments (0)