TL;DR: Treat every API key as three linked decisions: principal says which workload owns it, permission limits what that workload may do, and expiry limits how long the decision remains valid. For an automated prepaid-balance top-up, issue a separate credential to one replenishment workload, allow only the balance read and capped top-up operations it needs, and give it an explicit replacement deadline. The right design minimizes the blast radius of one leaked credential without making renewal so fragile that the balance runs dry.
That framing is more useful than calling a key "a password for an API." A password analogy hides the authorization policy and the clock. Those are exactly the parts that determine whether a stolen top-up credential can read unrelated accounts, move an uncapped amount, or remain useful long after its owner has changed.
What Do API Key Identity, Scope, and Lifetime Explain?
Identity answers "who is calling?" For machine-to-machine traffic, the useful answer is a workload, not the engineer who happened to create the secret and not an entire company account. In this example, prepaid-replenisher-prod is one principal. A reporting job, a staging instance, and a developer laptop are different principals even if they call the same service.
This separation is operational, not cosmetic. If production and staging share a key, the credential's blast radius includes both environments. If ten workers share one identity, an audit record can establish that the shared identity acted, but it cannot distinguish which worker held the secret. A key identifier may be safe to record for correlation; the secret value is not. In a prepaid system, that ambiguity also slows containment: responders must treat every holder of the shared key as potentially affected, replace the credential everywhere, and keep the balance automation alive while they do it.
Keep it boring.
Scope answers "what may this principal do?" It should describe allowed operations and resources. The replenisher needs to observe one prepaid balance and request a top-up under a policy cap. It does not need profile editing, credential administration, arbitrary transfers, or access to every funding account.
Lifetime answers "until when do we trust this credential?" Record an issue time, an expiry time, a status, and enough ownership metadata to replace the key deliberately. Rotation is a transition between two credentials, not a ritual that changes a string. During that transition, the old and new keys may both need a short, controlled overlap so instances can drain without an outage.
Short-lived is not automatically good. A five-minute lifetime paired with a renewal path that fails closed during a control-plane interruption can stop unattended replenishment. An unbounded lifetime avoids that availability edge but leaves a leaked secret valid until someone notices and revokes it. The explicit trade-off is exposure time versus dependence on renewal. Pick the lifetime from recovery and renewal behavior, then test both.
Step 1: Model the boundary before issuing the secret
Start with a credential record that contains no plaintext secret. The record below makes the three decisions reviewable and keeps the funding constraint close to the authorization data.
from dataclasses import dataclass
from datetime import datetime
from decimal import Decimal
@dataclass(frozen=True)
class CredentialPolicy:
key_id: str
principal: str
environment: str
allowed_actions: frozenset[str]
resource_ids: frozenset[str]
max_top_up: Decimal
issued_at: datetime
expires_at: datetime
status: str
The secret belongs in a secrets-management system, while the application consumes it through a narrow runtime interface. OWASP recommends centralized management, least privilege, expiration, rotation, revocation, auditing, and avoiding secret logging. Those controls reinforce one another: a well-scoped credential still needs prompt revocation, and a frequently rotated credential is still dangerous if every workload can retrieve it. The concrete example uses one production account, a 500-unit top-up cap, and a 30-day lifetime; those are test data, not universal defaults. A team must derive its own cap from the funding policy and its lifetime from how long issuance, deployment, verification, and emergency revocation actually take.
Do not put the account cap only in a runbook. Enforce it where the request is authorized. A useful decision rule is: if possession of one key permits a larger action than the unattended job is allowed to take, the scope is too broad.
Step 2: Enforce scope and lifetime on every call
Policy stored beside a key has no value unless the service checks it. The authorization function should reject an inactive or expired credential, an unexpected environment, an unapproved operation, the wrong prepaid account, and an excessive amount.
from datetime import datetime, timezone
from decimal import Decimal
class AuthorizationError(Exception):
pass
def authorize_top_up(
policy: CredentialPolicy,
*,
account_id: str,
amount: Decimal,
environment: str,
now: datetime | None = None,
) -> None:
checked_at = now or datetime.now(timezone.utc)
if policy.status != "active":
raise AuthorizationError("credential is not active")
if checked_at >= policy.expires_at:
raise AuthorizationError("credential has expired")
if environment != policy.environment:
raise AuthorizationError("environment is outside credential scope")
if "balance:top_up" not in policy.allowed_actions:
raise AuthorizationError("top-up action is outside credential scope")
if account_id not in policy.resource_ids:
raise AuthorizationError("account is outside credential scope")
if amount <= Decimal("0") or amount > policy.max_top_up:
raise AuthorizationError("amount is outside credential scope")
Keep authentication and business idempotency separate. A valid key proves that a request may be attempted; it does not prove that a retry should create a second top-up. The caller should attach a stable operation identifier to retries, and the receiving service should make duplicate handling explicit. This matters when a timeout leaves the caller uncertain whether the first request completed.
Log the key identifier, principal, policy decision, account alias, operation identifier, and reason code. Never log the supplied secret. Also avoid placing secrets in URLs, exception strings, or debug snapshots; OWASP calls out logs and other exposed locations as common places where secrets leak.
One sharp edge deserves a test: timestamps must be timezone-aware. Comparing a local naive timestamp with a UTC expiry can make a credential appear valid longer than intended or fail early around deployment boundaries.
One second matters.
Step 3: Rotate without creating a delivery gap
Credential replacement resembles a careful message-delivery rollout. Create the successor under the same reviewed policy, distribute it only to the intended workload, confirm that new instances authenticate with its identifier, and then revoke the predecessor. Do not revoke first. That order turns a routine security control into a replenishment outage.
A compact state machine is enough:
-
pending: created but not accepted for top-ups. -
active: accepted within its scope and lifetime. -
retiring: accepted only during the bounded overlap. -
revoked: rejected immediately, regardless of the recorded expiry.
Expiry and revocation solve different problems. Expiry establishes the latest normal acceptance time. Revocation is the emergency brake for suspected disclosure, ownership change, or a retired workload. Both checks belong in the request path. The limitation of a bounded overlap is that two credentials work briefly, increasing the number of valid secrets during migration; eliminating overlap shifts that risk into a hard cutover that can interrupt top-ups. Prefer overlap only when its duration, owners, and revocation condition are explicit.
Test the transition with a fake clock instead of waiting for wall time:
from datetime import datetime, timedelta, timezone
from decimal import Decimal
issued = datetime(2026, 9, 21, 8, 0, tzinfo=timezone.utc)
policy = CredentialPolicy(
key_id="key_replenisher_042",
principal="prepaid-replenisher-prod",
environment="production",
allowed_actions=frozenset({"balance:read", "balance:top_up"}),
resource_ids=frozenset({"prepaid_account_17"}),
max_top_up=Decimal("500.00"),
issued_at=issued,
expires_at=issued + timedelta(days=30),
status="active",
)
authorize_top_up(
policy,
account_id="prepaid_account_17",
amount=Decimal("125.00"),
environment="production",
now=issued + timedelta(days=29),
)
Then add negative tests for one second at expiry, a staging environment, another account, a revoked status, and an amount over the cap. Those failures are the contract. A happy-path test alone says little about blast radius.
Step 4: Operate the credential as a finite inventory
An inventory should answer four questions without revealing secrets: which principal owns each key, what it can reach, when it expires, and who responds if it is exposed. Alert early enough for a normal replacement and again when the remaining window threatens the rollout time. The exact interval depends on how quickly the team can issue, deploy, verify, and revoke; copying a universal number would hide that dependency.
Monitor authorization denials by reason, successful use by key identifier, use from an unexpected environment, and activity by a retiring key. A retiring key that remains busy is evidence that an old instance or forgotten consumer still exists. Investigate it before the overlap closes.
The audit trail needs restraint. Store policy changes and decisions, not credential values. Restrict access to that trail too; metadata connecting principals, accounts, and operational timing can be sensitive even when it contains no secret.
For incident readiness, rehearse one-key containment: revoke the affected identifier, verify that other principals continue working, issue a replacement with the same or narrower policy, and inspect usage associated with the old identifier. If revoking one replenisher key breaks reporting, staging, or credential administration, the identity boundary was already too large.
A compact rollout checklist
Migrate shared credentials one workload at a time. First inventory current consumers from approved metadata and audit events. Next define a dedicated principal and a resource-and-action policy for the prepaid replenisher. Issue its new key, deploy it through the secret store, observe successful use by the new identifier, and revoke the shared predecessor only after every expected instance has moved.
Finally, simulate expiry, emergency revocation, duplicate top-up retries, and loss of the renewal path. The design is finished when one credential can fail or leak without granting unrelated authority and without silently emptying the prepaid balance.
Sources
- OWASP, "Secrets Management Cheat Sheet": https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
Top comments (0)