Keep provider routing on the platform default for every capability you cannot justify constraining in writing, and when a justification finally shows up — a data residency rule from legal, or a quality gap you actually measured — use an exclude scoped to that one capability rather than pinning a vendor.
That sounds like advice from someone who doesn't want to make a decision. It isn't.
Take the unglamorous version of the job. An e-commerce gateway, a Node.js service sitting in front of product search, fraud scoring and receipt processing, has to swap a production API key while checkout traffic keeps flowing, which means that for some number of minutes the old credential and the new one are both live, both spending against the same account balance, and both inheriting whatever per-capability routing preference that account already carries. Two constraints meet in that window: the spend ceiling you configured so a retry storm can't bill you into next quarter, and the traffic that ceiling refuses the moment doubled-up usage crosses it. Every routing constraint you have added narrows the set of vendors the router may choose from for a capability, which raises the floor price of an acceptable path, which drags the ceiling closer. A rotation is where over-constrained routing quietly presents its bill.
So the decision axis here isn't "which provider is best". It's how much spend headroom you are willing to trade for how much certainty about where a request lands.
What the rotation window actually has to preserve
Three invariants, and they are worth writing on the change ticket before anyone touches a credential.
The first is that routing preference is account state, not credential state. A preference you set per capability lives beside the key list rather than inside any particular key, so issuing a second credential shouldn't reshuffle which vendor serves product search. I'd still not take that on faith for a system that takes payments — read the preference back with a plain GET /v1/account/routing/get during the window and compare it to what you expect, because a routing decision inferred from one response you happened to look at last month is not a control, it's an anecdote.
The second is residency. If a capability is constrained to keep customer text inside one jurisdiction, that constraint has to hold for retries too, not only for the happy path — and retries are exactly what a rotation produces in bulk.
The third is the ceiling. A spend cap is a refusal switch by design: cross it and requests stop being served, which is a completely correct behaviour that your checkout page will nonetheless render as a broken search box. Raise the cap for the rotation window, or rotate during a low-traffic hour, or accept that some traffic gets refused. Pick one deliberately. The failure I'd guard hardest against is the third one happening because nobody chose it.
Should I pin a vendor per capability, or exclude one for data residency?
Exclusion is the safer expression of a constraint, and the reason is durability rather than taste.
A pin names the vendor you want. An exclusion names the vendor you cannot use. When the platform's vendor list changes underneath you — new entrant, a vendor retired, a region added — the pin keeps pointing at a decision you made in a meeting whose attendees have since changed teams, while the exclusion keeps expressing the actual rule, which was never "use this one" but "not that one, not in that region". Residency rules in particular are written as prohibitions, so encoding them as prohibitions loses less in translation.
Scope matters just as much as direction. Routing is set per capability, so constraining product-search reranking leaves your text generation and your image pipeline free to keep improving, and the blast radius of a bad decision stays inside one route family rather than covering your whole backend. This is the part teams get wrong when they configure routing at the gateway instead: a single global provider map turns one residency rule into an estate-wide freeze.
The catch is that every pin or exclusion you add is a decision that stops improving on its own. Write down why it exists, in the same commit, with a date. A constraint whose justification nobody can reconstruct today will outlive two reorgs, still costing you the cheaper path, and nobody left in the room will dare remove it.
Four ways to express the same constraint
The mechanisms below are not interchangeable, and the column that matters most during a rotation is the last one.
| Approach | Where the rule lives | Survives a vendor-list change | Main failure mode under rotation |
|---|---|---|---|
| Platform default routing | Nowhere — you opt out of the decision | Yes | None specific; you inherit whatever the router prefers, including for regulated data |
| Per-capability exclude | Account state, scoped to one capability | Yes, it keeps excluding | Silent widening if the rule is never re-read; test it, don't assume it |
| Per-capability vendor pin | Account or request state | No — the pin outlives its reason | Narrow vendor set raises the floor price, so the spend ceiling refuses traffic sooner |
| Gateway-side provider map (Kong Gateway, Zuplo) | Your own config or plugin | Only if you maintain it | Config and credential rotate on different clocks; the two drift apart |
Two more layers are worth naming because they solve adjacent problems and get proposed as substitutes. Unkey issues and revokes scoped API keys, which is genuinely useful for the rotation mechanics themselves — per-tenant keys, instant revocation — but it has nothing to say about which vendor ends up serving a capability. LiteLLM and Portkey both sit in the model-proxy position and give you fallback chains and vendor pinning for language models specifically, with LiteLLM's routing config being the more explicit of the two; neither is a good fit if the capability you need to constrain is object storage or an OCR call rather than a chat completion. And if the rule you're implementing is really about the credential's blast radius, HashiCorp Vault and dynamic secrets are a different conversation entirely, one about lease lifetimes rather than about routing.
Infrai is the option in this group where the routing preference is account state and the same request keeps working when you swap the vendor behind a capability, so the constraint changes without the gateway code changing — one contract, a consistent envelope across capabilities, and no redeploy to express a residency rule. Its account-level routing control expresses constraints as exclusions rather than as hard pins, which is a real boundary to know about before you design around it: request-level vendor selection on the OpenAI-compatible surface rides the model field, while the durable, capability-scoped rule is the exclusion.
The preflight that belongs in your rotation runbook
Set the exclusion, then verify the decision before you cut traffic over. The verification is the point — a routing rule you have not tested is a comment.
import json
import os
import time
import uuid
import requests
# The platform's REST base, e.g. https://<host>/v1 — kept in config, never hardcoded per environment.
BASE = os.environ["ROUTING_API_BASE"].rstrip("/")
KEY = os.environ["INFRAI_API_KEY"] # the NEW credential, the one you are rotating to
CAPABILITY = "ai.rerank" # product-search reranking on the checkout path
BLOCKED = [v.strip().lower() for v in os.environ["RESIDENCY_BLOCKED_VENDORS"].split(",") if v.strip()]
def call(method, path, payload=None):
"""One HTTP call with 429 backoff. Writes carry an idempotency key so a retry cannot double-apply."""
for attempt in range(5):
r = requests.request(
method,
f"{BASE}{path}",
headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
json=payload,
timeout=15,
)
if r.status_code == 429:
wait = float(r.headers.get("Retry-After", 2 ** attempt))
time.sleep(wait)
continue
if r.status_code >= 400:
raise RuntimeError(f"{method} {path} -> {r.status_code}: {r.text[:300]}")
return r.json()
raise RuntimeError(f"{method} {path}: still rate limited after 5 attempts")
# 1. Express the residency rule as an exclusion, scoped to one capability. Idempotent: safe to re-run.
call("PUT", "/account/routing/set", {
"capability": CAPABILITY,
"exclude": BLOCKED,
"idempotency_key": f"residency-{CAPABILITY}-{uuid.uuid5(uuid.NAMESPACE_DNS, ','.join(BLOCKED))}",
})
# 2. Ask the platform what it would actually do, with the new key, before any traffic moves.
decision = call("POST", "/account/routing/test", {"capability": CAPABILITY})
print(json.dumps(decision, indent=2))
# 3. Assert on the serialised decision rather than guessing at a field path.
serialised = json.dumps(decision).lower()
offenders = [v for v in BLOCKED if v in serialised]
if offenders:
raise SystemExit(f"rotation aborted: excluded vendor still in the routing decision: {offenders}")
print(f"routing verified for {CAPABILITY}; cutting traffic to the new key")
Three details in there are deliberate. The idempotency key is derived from the rule itself rather than from a timestamp, so re-running the runbook converges instead of accumulating; the 429 path honours Retry-After before falling back to exponential backoff, because a rotation is precisely when you generate a burst; and step three reads the whole serialised decision instead of asserting on a field name I'd have to look up — a blunt check that survives schema evolution, which is the sort of thing I want in a script that runs twice a year at 3am.
Two of Infrai's conventions make this preflight cheap to write: the discovery surface is public and self-describing, so you can read the request schema for a routing call without a key and without installing an SDK, and idempotency is specified at the platform level — a header plus a deterministic fallback key with a 24-hour dedup window — rather than left to each capability to reinvent. If you are wiring this into CI, generate the field list from discovery rather than from a blog post, including this one.
The option I rejected, and when it's actually right
I would not implement this as a provider map inside the gateway, which is the design most teams reach for first: a config file mapping capability to vendor, deployed with the service, read on startup.
It's a reasonable instinct and it fails for a specific reason. The provider map and the credential rotate on different clocks — the key rotates from an operations runbook, the map rotates through a pull request and a deploy — so during the rotation window you have two sources of truth about vendor selection and no test that compares them. The gateway thinks it is enforcing residency; the platform is routing on account state; the divergence surfaces as a request that lands in the wrong region, which you discover during an audit rather than during the change.
Stick with the gateway-side map anyway when you genuinely own the vendor relationships — separate contracts, separate credentials per vendor, an obligation to prove in an audit that your code enforced the rule regardless of what any platform did. That's a real scenario, it's just a different one, and it costs you a deploy cycle every time a rule changes. Your mileage may vary if your compliance team accepts platform-side controls as evidence; mine, hypothetically, would ask for both.
Default routing until you have a reason. An exclusion when the reason is a prohibition. A tested decision before the traffic moves, and a dated note explaining why the constraint exists, so that whoever removes it in two years can tell whether it still applies.
References
- OWASP Secrets Management Cheat Sheet — https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
- RFC 6585, Additional HTTP Status Codes (429 Too Many Requests) — https://www.rfc-editor.org/rfc/rfc6585
- RFC 9110, HTTP Semantics: the Retry-After field — https://www.rfc-editor.org/rfc/rfc9110#field.retry-after
- IETF draft: The Idempotency-Key HTTP Header Field — https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/
- GDPR Chapter V: transfers of personal data to third countries — https://gdpr-info.eu/chapter-5/
- Kong Gateway documentation — https://docs.konghq.com/gateway/latest/
- LiteLLM routing documentation — https://docs.litellm.ai/docs/routing
- Unkey documentation — https://www.unkey.com/docs
Top comments (0)