DEV Community

YatesHolloway6872
YatesHolloway6872

Posted on

Zero downtime API key rotation under a hard spend ceiling and two alert thresholds

A small team should use all three controls, just not as a menu to pick from: threshold alerts tell you the shape of your spend, a scheduled budget review sets the number, and a hard stop keeps one bad night from eating the month's revenue. Ordering matters more than tooling. Set the hard stop well above your worst legitimate day, put the alert thresholds at fractions of it, and move the ceiling itself only on a fixed review date.

The part nobody warns you about is what happens to those controls when you rotate the credential they're attached to.

My service is a backend for a mobile game — leaderboards, a bit of matchmaking, some generated content — and it calls paid third-party APIs on the hot path. One person, weekly releases, no finance function. Every hour I spend on billing plumbing is an hour not spent on the game. So the budget question is never "what's the most rigorous control", it's "what's the cheapest control that still stops the thing that would actually hurt me".

Should a small team pick threshold alerts, a scheduled budget review, or a hard stop?

They aren't competing options, because they answer different questions.

Control Question it answers What it costs when it fires
Threshold alerts Is spend drifting away from plan? A notification you might sleep through
Scheduled budget review Is the ceiling still the right number? Half an hour, once a month
Hard stop Am I about to lose more than this feature earns? Refused traffic, which is the expensive kind

The axis that decides how aggressive you get is spend ceiling versus refused traffic. For a game backend those two aren't symmetrical at all. An overspend of forty dollars is forty dollars. A failed leaderboard write during a weekend event turns into support tickets, a thread on the game's subreddit, and a one-star review that outlives the incident by a year.

That asymmetry is the whole design. My ceiling sits at roughly three times a busy day, my alert thresholds sit at 60% and 85% of the ceiling, and the hard stop doesn't refuse everything when it trips — it sheds by traffic class. Cosmetic generated content goes first. Purchase validation never goes.

Scheduled review is the control people skip because it produces no artifact. It's also the only one that catches slow drift: a feature you shipped in March that quietly tripled its call volume by September won't set off an alert tied to a ceiling you raised twice along the way.

The rotation window is where spend controls actually break

Every budget control you can buy or build is keyed to something: an API key, a project, an organization. Rotation moves traffic between two of those things, which means rotation is exactly when your counters lie to you.

Two failure modes, both quiet. If your ceiling is enforced per credential — a spending limit configured on the key itself — then during an overlap window you have two live keys each carrying the full limit, and your effective ceiling has silently doubled. Nothing errors. Nothing alerts. You find out on the invoice. The mirror image is just as bad: if your counter is derived per key from the provider's usage view, the retiring credential's charges land in a different bucket than the new one, so the aggregate you alert on reads low during the precise window when you are least sure your new key is wired correctly. I assumed for a long time that the provider's dashboard was the source of truth here. It isn't, or at least it isn't shaped the way a budget control needs — it's shaped per credential, and my budget isn't.

The fix is boring and it's the reason the rest of this piece has any structure: make the budget bucket a property of the service, not of the credential. One bucket, game-api, and every charge recorded against it carries the credential id as a label rather than as its identity. Then rotation is a label change, not a counter reset.

With that in place the rotation itself is the standard overlap dance the OWASP secrets management guidance describes: mint the new credential, deploy it as the secondary, promote it to primary, keep the old one valid long enough that every instance and every queued retry has moved, then revoke. I keep the overlap at 24 hours because my longest job retry chain is under an hour and 24 hours is one sleep cycle plus margin. Zero downtime comes from the overlap, not from clever code.

app-secrets set PROVIDER_KEY_SECONDARY "$NEW_KEY"    # old credential still valid
app-secrets promote PROVIDER_KEY_SECONDARY           # new key carries live traffic
provider-admin keys revoke "$OLD_KEY_ID"             # 24h later, ledger shows no charges on it
Enter fullscreen mode Exit fullscreen mode

Secret stores like HashiCorp Vault and AWS Secrets Manager will mint and distribute that credential on a schedule, but they don't know what it costs — rotation and spend are separate concerns in every one of them. Gateways such as Kong Gateway or Tyk can cap requests per key, but requests aren't dollars: one expensive model call and one trivial lookup are indistinguishable to a request counter. Usage-metering services like OpenMeter and Amberflo do the dollar math properly, at the price of another integration to own.

The smallest thing that worked

Two hundred lines, one table, no new dependency. The ledger records what each call cost, the gate reads today's total before the next call, and both are rotation-agnostic because they key on the bucket.

type Credential = { id: string; secret: string; state: "primary" | "retiring" };

const BUCKET = "game-api";
const CEILING_CENTS = 21_000;          // ~3x a busy day
const ALERT_AT = [0.6, 0.85];          // fractions of the ceiling

type Charge = { bucket: string; credentialId: string; cents: number; at: number };

// In production this is one Postgres table with an index on (bucket, at).
const ledger: Charge[] = [];

function spentToday(bucket: string): number {
  const dayStart = Date.UTC(new Date().getUTCFullYear(), new Date().getUTCMonth(), new Date().getUTCDate());
  return ledger
    .filter((c) => c.bucket === bucket && c.at >= dayStart)
    .reduce((sum, c) => sum + c.cents, 0);
}
Enter fullscreen mode Exit fullscreen mode

The gate is where the refused-traffic tradeoff gets encoded, and it's the only opinionated part. Each traffic class gets its own multiple of the ceiling, so the ceiling behaves like a budget for cosmetic work and like a fuse only for the stuff I'd rather overspend on than drop.

type TrafficClass = "revenue" | "retention" | "cosmetic";

const SHED_ABOVE: Record<TrafficClass, number> = {
  cosmetic: 0.85,   // first to go
  retention: 1.0,
  revenue: 1.15,    // I will happily blow 15% past the ceiling to validate a purchase
};

function admit(cls: TrafficClass): boolean {
  return spentToday(BUCKET) < CEILING_CENTS * SHED_ABOVE[cls];
}

async function paidCall(cls: TrafficClass, body: unknown, creds: Credential[]) {
  if (!admit(cls)) throw new Error(`budget_shed:${cls}`);
  const cred = creds.find((c) => c.state === "primary")!;
  const res = await fetch(`${process.env.PROVIDER_BASE}/completions`, {
    method: "POST",
    headers: { authorization: `Bearer ${cred.secret}`, "content-type": "application/json" },
    body: JSON.stringify(body),
  });
  if (res.status === 429) return retryAfter(res);         // Retry-After, then the same gate again
  const json = await res.json();
  ledger.push({ bucket: BUCKET, credentialId: cred.id, cents: estimateCents(json.usage), at: Date.now() });
  return json;
}
Enter fullscreen mode Exit fullscreen mode

Worth knowing for the client side of this: HTTP has a status code that means exactly "you're out of budget", 402 Payment Required, and RFC 9110 still lists it as reserved. So nobody reliably sends it and nobody handles it. In practice budget exhaustion arrives as a 429 or a 4xx with a vendor-specific body, which is why the gate above refuses locally rather than waiting to be told.

What I would change at ten times the traffic

The in-process array goes first, obviously. After that, three things in order: reconcile the ledger against the actual invoice monthly, because a token-based cost estimate drifts and an unreconciled estimate slowly becomes fiction; emit today's spend as an OpenTelemetry gauge so the threshold alerts live next to every other alert instead of in a bespoke email job; and move the shed decision to one place at the edge rather than at each call site, because a scattered gate is a gate somebody forgets.

I'd also stop rotating on a calendar and rotate on events — a contractor leaving, a leaked log line, a dependency compromise. Calendar rotation is a compliance ritual for a one-person shop. Event rotation is the thing that would actually save me, and it only works if rotation is so cheap I'll do it at 2am, which loops back to the overlap window and the bucket-not-credential rule.

Where this is the wrong call

A hard stop is wrong wherever refused traffic costs more than the overspend it prevents. Payment validation, anti-cheat, auth — for those paths, stick with alerts plus whatever spending limit the provider offers as a last resort, and let the invoice be the surprise. I'm not sure there's a principled way to pick the revenue multiple; 1.15 is a number I can defend for my margins and not for yours.

If you have a finance function and a procurement cycle, the scheduled review is your real control, and a hard stop just manufactures incidents that someone has to page you about. The catch with everything above is that it assumes usage-based billing on the hot path. On flat monthly pricing, none of it applies — you have a contract, not a ceiling.

And if your API spend is a rounding error against payroll, don't build any of this in 2026. Set one alert, put a recurring 30-minute review in the calendar, and go ship the game.

Further reading

Top comments (0)