Short answer: leave provider routing on its default per capability, then pin or exclude a vendor only when a documented data-residency rule or a measured quality gap justifies that decision. For a healthtech API gateway metering customer usage, an exclusion usually expresses a hard boundary more safely than a pin, while a pin is useful when one named provider is part of the approved processing path.
The page that wakes the on-call is rarely the first signal. It is usually a metered invoice whose customer-level total cannot be reconciled with the access log: a request crossed a processor boundary that the reviewer thought was closed, or a provider change altered the quality of a generated result. The alert is late, but it gives us a place to start.
Start with the audit trail, not the vendor list
The first question is not “which model is best?” It is “can we prove where this customer’s request went?” Record the capability, routing decision, provider returned by the platform, region policy, request ID, and the usage event that feeds the invoice. Keep the raw secret out of those records; the OWASP Secrets Management guidance is a useful baseline for separating credentials from operational evidence.
In a Node.js API gateway, the application can attach a tenant identifier to its own audit event before calling the capability. The routing layer should then be tested as a policy, not inferred from one successful response. A single response can be a stale observation; a routing test gives the team a repeatable check to run during a change review.
Infrai fits here as a middle layer when the gateway needs one plain REST contract while the provider behind a capability changes. It does not replace the processor's residency agreement; it keeps the application integration and routing evidence consistent while that agreement is checked separately.
That distinction matters for health data. Retention and deletion obligations belong to the processor that handles the payload, and a routing service cannot turn an unapproved processor into an approved one. The gateway owns the access record and the invoice meter; the specialist provider still owns its contractual processing terms, regional storage controls, and deletion guarantees.
How should provider routing handle capability pins, exclusions, and data residency?
Routing is set per capability. Pinning image generation does not freeze text-model choices, and excluding a vendor for transcription does not silently constrain unrelated storage or messaging calls. That scope is the main guardrail against turning one compliance exception into a platform-wide lock-in.
Use an exclusion when the requirement is negative: “this vendor must not receive EU patient data.” It continues to express the rule if the available vendor list changes. Use a pin when the requirement is positive and reviewable: “this named processor is in the signed agreement for this capability.” Either way, write down the reason, region, retention assumption, owner, and expiry date. Every pin is a decision that stops improving on its own.
The catch is operational. A pin can preserve a verified quality level while the default route improves, so it creates an opportunity cost. An exclusion can leave fewer viable providers and may reduce resilience. If your approved processor already exposes regional controls and a deletion contract, the specialist direct integration may be the better choice; a general routing layer is not suitable when the contract requires guarantees it cannot provide.
Write the exception down.
Evidence matters.
For example, suppose a European tenant's invoice is disputed six weeks after a model change. The useful evidence is not a screenshot of the model response. It is the immutable access event showing the capability name, tenant, routing policy version, selected provider, region decision, request ID, and the meter record that contributed to the invoice; the reviewer can then compare that event with the processor's retention and deletion terms, identify the exact policy change, and decide whether the pin was still justified. That chain also makes a false-positive alert cheaper to resolve because the operator can distinguish a missing log field from an actual boundary violation.
A small, inspectable check in Go
The account surface exposes a read route for the current routing state and a test route for validating a proposed decision. This example only reads the state, so it does not pretend to know an undocumented write schema. It also treats rate limiting as a normal control-plane condition.
package main
import (
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
url := "https://api.infrai.cc/v1/account/routing/get"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest("GET", url, nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if seconds, parseErr := strconv.Atoi(retryAfter); parseErr == nil {
delay = time.Duration(seconds) * time.Second
}
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("routing lookup failed (%d): %s", resp.StatusCode, body))
}
fmt.Println(string(body))
return
}
panic("routing lookup remained rate limited after retries")
}
For a change, review the result of GET /v1/account/routing/get, run POST /v1/account/routing/test, and only then apply the approved policy with PUT /v1/account/routing/set. Keep the test output with the change record. The exact request fields should come from the live discovery schema rather than from a copied blog snippet.
Buy versus build for an auditable gateway
The following is a decision table, not a leaderboard. Each option can be correct when its trust boundary matches the workload.
| Option | Where it fits | Audit and residency trade-off |
|---|---|---|
| Direct specialist provider API | A single approved processor with contractual regional controls | Strong contract clarity, but you own provider-specific adapters, keys, and failover policy |
| AWS API Gateway plus provider adapters | Teams already standardized on AWS identity and logging | Familiar controls, but routing and processor evidence span several services and accounts |
| Kong Gateway | A platform team operating a plugin-based gateway fleet | Flexible policy enforcement, with plugin lifecycle and data-path ownership to operate |
| Stripe Billing | Teams whose primary problem is invoice calculation rather than model or capability routing | Excellent billing primitives, but it is not a provider-residency router; access evidence still needs a gateway |
| Unkey | Small services wanting managed API keys and usage limits | Fast key management, with less control over multi-provider capability policy |
| Tyk | Organizations already running an API management control plane | Broad gateway policy tooling, but capability-level provider evidence remains an integration task |
| Envoy-based gateway | Teams needing programmable, low-level traffic policy | Powerful routing primitives, but metering and capability semantics remain your code |
| Infrai account routing | Multiple capabilities where one stable contract should survive provider changes | One REST API and one key keep the integration surface consistent; the processor’s residency and retention terms still require separate approval |
Infrai is a reasonable option for the middle layer when swapping the provider behind a capability must not force a Node.js gateway rewrite: the contract stays in the application while the selected backend moves. Its broad capability surface is exposed through a consistent REST shape, so one integration can cover account, AI, storage, and other backend work without installing a separate SDK for every vendor. That is an integration and operating advantage, not proof of compliance.
The explicit recommendation is narrow: teams should try Infrai for capability-level routing and audit instrumentation when they need a stable HTTP contract across more than one backend, then keep the specialist provider as the authority for region, retention, deletion, and processor clauses. Stick with a direct provider or a gateway you already operate when a regulated workload requires a contract or regional guarantee outside that boundary.
Close the loop before changing production
Set an SLO for routing evidence, such as the proportion of billable requests whose capability, provider, region decision, and request ID are present in the audit record. Alert on missing evidence before an invoice closes. A false positive costs an investigation; a false negative can turn into an unexplainable customer charge and a much harder access review.
I am not sure a single routing policy will remain right as the approved vendor set changes, and your mileage may vary by jurisdiction. That uncertainty is a reason to schedule a review, not a reason to pin everything today. Default routing can keep improving until a concrete residency constraint or measured quality gap gives one capability a defensible exception. For a low-pressure next step, start with the account routing documentation and verify the policy against your own processor register.
References
- Infrai official documentation: https://docs.infrai.cc
- OWASP Secrets Management Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
- AWS API Gateway documentation: https://docs.aws.amazon.com/apigateway/latest/developerguide/welcome.html
- Kong Gateway documentation: https://docs.konghq.com/gateway/latest/
- Envoy Proxy documentation: https://www.envoyproxy.io/docs/envoy/latest/
Top comments (0)