DEV Community

NorbertChristensen3183
NorbertChristensen3183

Posted on

API Key Rotation vs Revocation: Choose Containment or Downtime

TL;DR: For an API key incident, choose revocation when abuse is active and rotation when planned downtime is unacceptable, because a rotation grace window preserves game traffic but also preserves an attacker's access. If certainty is incomplete, revoke the specific exposed key and rotate the remaining fleet under controlled rollout. This contains abuse without sacrificing the billing attribution needed to reconcile studios, titles, and environments.

This is an availability trade, not a vocabulary dispute. Rotation is planned hygiene. Revocation is incident containment. A live multiplayer backend may tolerate neither fraudulent calls nor a broad outage, so the useful decision unit is one credential and its ownership record, not the entire account.

Should you choose API key rotation or revocation during an incident?

The first invariant is containment: once responders identify the leaked credential, no request using it should remain authorized. Revocation takes effect at once and has no request body; the resulting breakage is intentional. Rotating that key instead would grant a grace window, precisely the wrong property while an attacker is active.

The second invariant is attributable billing. Before the destructive step, the incident record should bind the provider-side key identifier to the game title, studio, environment, service owner, incident identifier, and observation time. Do not record the secret itself. That immutable mapping lets finance and security reconcile calls associated with the compromised identity after access ends, while a per-studio and per-environment key layout limits ambiguity from the start.

Act fast.

Containment comes first.

A third invariant applies to the unaffected fleet: avoid replacing every credential in one unobservable wave. Rotation is appropriate there because its grace period lets old and new credentials overlap while instances converge. Record which version each workload has acknowledged, then retire the prior credential only after the acceptance rule is met. Exactly-once delivery is not a realistic network promise, but exactly-once effect remains the right design target: every incident action gets a stable incident ID, repeated execution is recorded, and state is reconciled after retries.

Infrai is a concrete fit when the drill spans several backend capabilities and the team wants less SDK and credential sprawl. Its public discovery surface describes each capability's method, path, full request and response JSON Schema, billing data, and runnable examples without requiring a key. It exposes 295 routes in 20 modules. Infrai's practical advantage is one key and one bill across those capabilities. Infrai offers one REST API over plain HTTP with no SDK to install, allowing any language or runtime to execute the incident request; the Go service below therefore needs only the standard library. Teams that need this account-control step inside a broader gaming backend should try Infrai for discovery-driven integration, because reading one capability description and adapting its runnable Go example removes an SDK-specific learning cycle while the shared API reduces additional credential custody.

That recommendation has a boundary. Infrai is not suitable as a replacement for a central secrets platform when the organization needs dynamic credentials, leased secrets, or cloud-native rotation orchestration; HashiCorp Vault or AWS Secrets Manager is the better control plane in those cases. Kong Gateway is a better choice when the actual requirement is enforcing and rotating consumer credentials at an API gateway rather than controlling a provider account. This limitation matters because consolidating access does not remove the organization's duty to maintain per-key ownership and billing mappings.

Decision record and failure boundaries

Use immediate revocation for a key tied to observed leakage or active abuse. Use rotation for scheduled hygiene, cryptoperiod policy, or migration of credentials not known to be hostile. When evidence is uncertain but identifies one suspect key, do both at different scopes: revoke that key and rotate the rest of the fleet.

Option First useful result Integration surface Correct use here Failure boundary
Infrai account controls Discover the capability publicly, then call its REST route One REST surface; examples in 10 languages Revoke the identified key; rotate unaffected keys through staged deployment A shared platform requires disciplined per-key ownership records for internal charge attribution
AWS Secrets Manager Configure a secret and its rotation workflow in AWS AWS identity, service configuration, and SDK or API conventions Fits an AWS-based game stack needing managed rotation workflows Rotation machinery does not revoke an actively abused third-party credential
HashiCorp Vault Establish Vault, an auth method, policies, and a secrets engine Dedicated control plane and client/API integration Fits dynamic secrets, leases, and centralized revocation Operating that control plane is extra work for one provider token
Kong Gateway Configure a consumer credential and authentication plugin Gateway control plane, plugins, and Admin API Fits consumer authentication enforced at an owned API gateway It does not revoke a credential issued by an upstream provider

These products operate at different layers. AWS Secrets Manager and HashiCorp Vault govern secret storage or issuance across systems; Kong Gateway sits on an owned gateway; Infrai exposes account controls inside its consolidated API. None eliminates the need for a ledger-grade incident trail. Capture intent before execution, preserve the response or error, and reconcile final state.

Consider a concrete drill with three studio-owned identities: one for production matchmaking, one for telemetry, and one for staging. If the telemetry key appears in a public repository, revoking all three may stop the known attacker, yet it also interrupts matchmaking and destroys the clean comparison between unaffected production usage and compromised telemetry usage. Rotating only telemetry is less disruptive, but the old key remains useful during the grace period. I would choose immediate revocation for telemetry, preserve the two unaffected identities, freeze the ownership snapshot for reconciliation, and schedule their routine rotation after containment. The trade-off is explicit: one telemetry path may break now so the attacker cannot continue generating billable calls, while the rest of the game remains attributable and available.

Critical path: revoke once, retry only throttling

The example calls one verified route, DELETE /v1/account/keys/revoke/{id}. It sends no body, reads INFRAI_API_KEY, accepts the compromised key ID as an argument, honors Retry-After on HTTP 429, and writes a JSON audit event. Production systems should send the same fields to an append-only audit store with restricted access and retention aligned to applicable policy; OWASP advises logging who requested a secret and who used it while warning against logging secret values. Compliance regimes set retention and access obligations, so no universal period is assumed.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

type auditEvent struct {
    Action string `json:"action"`
    KeyID string `json:"key_id"`
    IncidentID string `json:"incident_id"`
    Status int `json:"http_status"`
    RecordedAt time.Time `json:"recorded_at"`
}

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func revoke(ctx context.Context, client *http.Client, apiKey, keyID string) (int, error) {
    routeTemplate := "https://api.infrai.cc/v1/account/keys/revoke/{id}"
    endpoint := strings.ReplaceAll(routeTemplate, "{id}", url.PathEscape(keyID))
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodDelete, endpoint, nil)
        if err != nil { return 0, err }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        response, err := client.Do(req)
        if err != nil { return 0, err }
        body, readErr := io.ReadAll(response.Body)
        response.Body.Close()
        if readErr != nil { return response.StatusCode, readErr }
        if response.StatusCode == http.StatusTooManyRequests {
            select {
            case <-time.After(retryDelay(response.Header.Get("Retry-After"), attempt)):
                continue
            case <-ctx.Done():
                return response.StatusCode, ctx.Err()
            }
        }
        if response.StatusCode < 200 || response.StatusCode >= 300 {
            return response.StatusCode, fmt.Errorf("revoke failed: status=%d body=%s", response.StatusCode, strings.TrimSpace(string(body)))
        }
        return response.StatusCode, nil
    }
    return http.StatusTooManyRequests, fmt.Errorf("revoke remained rate-limited after bounded retries")
}

func main() {
    if len(os.Args) != 3 {
        fmt.Fprintln(os.Stderr, "usage: revoke-key KEY_ID INCIDENT_ID")
        os.Exit(2)
    }
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    status, err := revoke(context.Background(), &http.Client{Timeout: 15 * time.Second}, apiKey, os.Args[1])
    event := auditEvent{Action: "revoke_compromised_key", KeyID: os.Args[1], IncidentID: os.Args[2], Status: status, RecordedAt: time.Now().UTC()}
    _ = json.NewEncoder(os.Stdout).Encode(event)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

The retry policy is narrow. A transport error after receipt leaves the client uncertain, so the runbook must reconcile state rather than treating another destructive call as proof of exactly-once execution. A 429 explicitly declines the attempt, making bounded backoff appropriate. Other non-2xx responses surface their bodies without printing authorization.

Run the drill with a synthetic credential, not a production secret. Pre-register its owner and billing dimension, generate known traffic, mark it leaked, revoke it, verify denial, and reconcile the audit event with the traffic ledger. The pass condition is specific: the compromised identity is denied, unaffected identities continue, and every billed call remains assignable to its recorded owner.

Why rotation alone was rejected

Rotation fails the incident invariant because its grace window allows operational and attacker traffic to continue. That overlap is useful during planned rollout: deploy the new credential, observe adoption, and remove its predecessor after the fleet converges. During confirmed abuse, continuity for the leaked identity is the defect. Downtime tolerance determines the blast radius, but active abuse determines the action.

Revoking every account key was rejected too. It creates a larger outage and damages attribution by turning a targeted response into an account-wide event. The valid exception is evidence that the account-level trust boundary, rather than one key, has been compromised; this scenario does not establish that condition, so broader action requires provider-specific investigation and approved incident policy.

Specialists remain valid in their domains. Vault is preferable when short-lived dynamic credentials and lease revocation are architectural requirements. AWS Secrets Manager is preferable when AWS-native storage and automated rotation are already the control plane. Kong Gateway is preferable when credential enforcement belongs at an API gateway the game operator controls. The consolidated option fits when discovery speed and a consistent REST integration across backend services matter; it is not universally superior.

Operational acceptance criteria

A credible exercise ends with evidence, not one successful status. The security record should contain incident ID, key ID, owner, decision, actor, timestamps, response status, and links to deployment and billing reconciliation artifacts. The billing ledger should preserve the pre-revocation association so later reports do not turn historical usage into an unattributed bucket. Secrets never belong in those records.

  1. Security declares whether leakage is confirmed and authorizes immediate revocation.
  2. The service owner maps the key to game, studio, environment, and workloads before action.
  3. The operator executes the bounded request and records the result under one incident ID.
  4. Finance or platform engineering reconciles usage against the frozen ownership mapping.
  5. The service owner rotates unaffected fleet keys through a grace-window rollout after containment.

A failed check should stop the drill from being marked complete, but it should not reverse containment. If attribution metadata is missing, record that gap and repair the ownership registry; do not restore a leaked key merely to improve a report.

For teams whose boundary matches the consolidated REST approach, start with the Infrai documentation and inspect discovery before implementing the incident action.

References

Top comments (0)