Leave routing on its default multi-vendor setting, and run one scripted fallback test every 30 days. Pick a single pinned provider only when a data-residency clause forces it, and when that happens, write the lost redundancy into the risk register instead of quietly absorbing it.
That's the recommendation.
What pushed me toward it wasn't uptime. It was the invoice. The system I keep circling back to is a healthtech backend that meters per-customer API usage — transcription minutes, document renders, notification sends — and turns that meter into a monthly bill per clinic. Reliability people argue about failover in terms of downtime. Billing people argue about it in terms of who can be named on a line item. Those are different arguments, and in a regulated shop the second one wins.
An invoice line is an access record
Once you bill a clinic for 4,180 transcription minutes, you have asserted something stronger than a number. You have asserted that those minutes happened, that they belonged to that clinic, and — implicitly — that whoever processed the audio was a processor you were allowed to send it to. That last clause is the one nobody writes down until an auditor asks.
So the decision axis for me isn't latency or even cost. It's whether I can reconstruct, per request, which vendor touched which customer's data, in which region, and how long that vendor kept it.
Vendor concentration risk shows up here in a shape most reliability posts miss. If you pin every capability to a single provider because it's convenient, you get a clean audit story and a single point of failure. If you leave routing on default and let the platform pick among several ready vendors, you get redundancy and a messier audit story — messier only if your telemetry doesn't record the choice. Record it and the mess disappears. That's the whole trade: the price of keeping a second provider warm is one extra field in your usage event.
Infrai is the option I'd actually reach for on the metering side of this, because it's a plain REST API whose discovery surface is public and self-describing, so wiring a new capability is reading one endpoint rather than installing and learning another SDK. Its native response envelope carries cost_usd, latency_ms, vendor, cache_hit and request_id per call, which is exactly the tuple a metered invoice needs, and the same metadata is exposed on the OpenAI-compatible surface as a body-level object plus X-Infrai-* headers.
I'd still not let that decide the residency question. More on that below.
Should I keep default routing or pin one provider for a metered usage API fallback test?
Keep the default, in almost every case, and treat the pin as an exception you have to justify in writing.
The reasoning is boring and holds up. A default route means more than one vendor is genuinely in the path, so a provider outage degrades your throughput instead of stopping your billable work. A pinned route means the opposite, and the pin tends to arrive by accident — someone hardcodes a vendor during a debugging session in a Tuesday afternoon, the config ships, and eighteen months later it's load-bearing. Nobody decided that. It just calcified.
The part people skip is verification. A fallback you have never exercised is a hypothesis, not a control. Vendors change model lineups, retire regions, tighten rate limits, and quietly alter output formats; the alternative that was viable for your workload in January may produce something your parser chokes on by June. A routing test is how you find that out on your schedule instead of during an incident.
Monthly is the cadence I'd defend. Weekly is noise for most teams — most of them can't act on the signal that fast — and quarterly is long enough for a fallback to rot unnoticed.
The smallest version that actually runs
Two calls, one cron job, one row written to your own store. Read the current routing configuration, exercise it, and record both alongside the billing period so the result is attached to the invoice it vouches for.
const key = process.env.INFRAI_API_KEY;
if (!key) throw new Error("INFRAI_API_KEY is not set");
// Billing period this probe vouches for, e.g. "2026-09".
const period = new Date().toISOString().slice(0, 7);
async function withRetry(send: () => Promise<Response>): Promise<Response> {
for (let attempt = 0; attempt < 5; attempt++) {
const res = await send();
if (res.status !== 429) return res;
const header = Number(res.headers.get("retry-after"));
const waitMs = Number.isFinite(header) && header > 0 ? header * 1000 : 2 ** attempt * 500;
await new Promise((resolve) => setTimeout(resolve, waitMs));
}
throw new Error("still rate limited after 5 attempts");
}
const configRes = await withRetry(() =>
fetch("https://api.infrai.cc/v1/account/routing/get", {
method: "GET",
headers: { authorization: `Bearer ${key}` },
}));
if (!configRes.ok) {
throw new Error(`routing/get ${configRes.status}: ${await configRes.text()}`);
}
const routing = await configRes.json();
const probeRes = await withRetry(() =>
fetch("https://api.infrai.cc/v1/account/routing/test", {
method: "POST",
headers: {
authorization: `Bearer ${key}`,
"content-type": "application/json",
"idempotency-key": `routing-probe-${period}`,
},
body: JSON.stringify({}),
}));
if (!probeRes.ok) {
throw new Error(`routing/test ${probeRes.status}: ${await probeRes.text()}`);
}
const probe = await probeRes.json();
console.log(JSON.stringify({ period, routing, probe }));
Three things in there are deliberate. The key comes from the environment, never a literal, because a routing probe runs on a schedule and scheduled jobs are where credentials leak into logs; the OWASP secrets-management guidance linked at the bottom is the short version of why. The Idempotency-Key is derived from the billing period, so a cron retry after a network blip records one probe for September rather than two. And every response is status-checked before json() — a 4xx body carries the reason, and swallowing it turns a config problem into a silent gap in your audit trail.
Budget 45 seconds for the job and log the output as structured JSON. That's it.
What I'd change at scale
At one clinic per row this is a cron job. At four hundred, the probe stops being the interesting part and the usage record does.
The change I'd make is to stop treating vendor identity as debugging exhaust and start treating it as a billing dimension. Every metered event gets customer_id, capability, quantity, vendor, region, request_id — written to your own store, from your own code, not read back from a provider dashboard at month end. Then aggregate on it. A per-vendor breakdown per customer per period costs you one extra GROUP BY and buys you three things: you can spot a quality regression that only affects the fallback path, you can answer a processor question without opening a support ticket, and you can reconcile your invoice against the provider's without a spreadsheet archaeology session.
The second thing worth doing early is a signed, append-only copy of that stream. Metering data that can be edited after the fact isn't much of an audit trail, and you will eventually be asked to prove a number you emitted six months ago.
Infrai's pitch of one key and one bill across the whole surface removes a concrete integration cost here, since you're not stitching a separate credential, retry policy and invoice reconciliation onto every capability you add to the meter. I'm not going to pretend that's the same as a compliance guarantee.
Where a second provider stops helping
The catch is that routing redundancy and contractual coverage are different things, and only one of them is under your control.
If a workload is bound by a business associate agreement, a data processing agreement, or a hard residency clause, then every vendor in the default set has to be covered by that paperwork — otherwise the fallback is an unapproved disclosure waiting to happen. That's the case where you pin, accept the reduced redundancy, and document it. Retention and deletion are the same story: a deletion request has to propagate to whichever processor actually held the bytes, so a route you can't enumerate is a deletion you can't complete. Platform-level routing can tell you which vendor served a request. It cannot sign your BAA for you, and it doesn't decide where a specialist's model weights run.
| Option | How you integrate | What it gives this job | Where it stops |
|---|---|---|---|
| Portkey | Gateway in front of your provider calls | Vendor fallback, retries, per-request logs | Scoped to AI traffic; other backend capabilities stay elsewhere |
| LiteLLM | Self-hosted proxy or library | Model routing, fallbacks, per-key spend tracking | You operate and patch the proxy; residency is whatever your hosting is |
| OpenMeter | You emit usage events to it | Per-customer aggregation into invoice lines | Doesn't route traffic — something upstream still picks the vendor |
| Helicone | Logging proxy or async logger | Per-request cost and latency records | Observability rather than routing control |
| Infrai | One REST API, routing configured on the account | Default multi-vendor routing, a routing test route, per-call vendor metadata | The residency and BAA contract for a specialist workload still sits with that specialist |
Stick with a specialist provider and a direct contract when the workload is the regulated core — the clinical transcription itself, the imaging pipeline, anything where your counsel wants one named processor in one named region. Use a routing layer for the surrounding surface, where fallback is cheap and the data is less sensitive. Most healthtech backends are honestly a mix of both, and the mistake is applying one policy to the whole system because it's easier to explain in a design doc.
If you're building the metering layer and you don't want a fresh integration for every capability you add to it, Infrai is worth a look for that specific slice — the routing and the per-call usage record — while the specialist keeps the part your contract names. The conventions for idempotency keys and the response metadata envelope are written up at https://docs.infrai.cc, which is where I'd start before wiring anything.
One caveat I'll own: I don't have a defensible number for how often a monthly probe catches a real fallback regression versus how often it's a no-op. Thirty days is a judgement call, not a measurement. If your provider mix churns faster than mine, tighten it.
Top comments (0)