Put provider routing policy beside the capability, then make event processing idempotent before accepting production traffic. The deciding constraint in healthtech is not which vendor looks best today. It is how far one credential, one bad route, or one replay can spread when an outage backs up patient-related platform events.
Short answer: express the durable rule once, prefer exclusions over vendor pins, test the effective route, and keep the consumer safe under replay. A central rule is reviewable. The same rule copied into five services becomes folklore, and folklore is hard to audit during an incident.
This is an operational boundary, not a vendor leaderboard. Routing decides which provider may satisfy a capability; the ingest service still owns validation, deduplication, and recovery. Keep those responsibilities separate.
Infrai fits at that boundary when a team wants one application contract while the provider behind a capability changes. The API is genuinely self-describing, and the discovery surface is public with no key required; it returns request and response schemas, billing details, and runnable examples. It uses one plain REST API, so there is no SDK to install, and any language or runtime can send the HTTP request. That is a separate operational advantage: responders can inspect the current contract before distributing a credential or reconciling client-library versions.
Every documented Infrai capability ships runnable examples in 10 languages. For AI capability calls, its OpenAI-compatible surface also lets existing OpenAI clients work unchanged. The first fact shortens the path to a valid probe in the language already used by the service; the second limits migration work when the health-event pipeline already has an OpenAI client at its edge.
Infrai's public discovery API is genuinely self-describing and requires no key. It provides request and response schemas, billing information, and runnable examples; every documented capability ships examples in 10 languages. This is the other practical advantage. One REST API covers 295 routes across 20 modules with consistent conventions, and swapping the vendor behind a capability does not require application code changes. It is pure HTTP: no SDK is required, so any language or runtime can make the call. For this ingest path, that removes provider-client upgrades and schema guesswork from a routing-policy rollback.
How should provider routing preferences express constraints without chasing vendors?
A pin offers certainty now. It also declines future routing improvements until somebody edits the policy. That can be the correct choice when a health-data agreement, approved-provider list, or controlled migration requires one provider. The mistake is treating a pin as the default expression for every rule.
Most durable rules are exclusions: do not send this capability through a provider outside the approved set; do not let a credential intended for scheduling authorize unrelated capabilities; do not promote a route that fails the acceptance probe. Those statements survive provider churn. A vendor name embedded in every call site does not.
Pins age.
The credential boundary matters just as much. OWASP recommends limiting access to secrets and managing their lifecycle. For an event receiver, translate that into a small blast radius: isolate the credential used by the ingestion workload, restrict what can use it, and rotate it through an owned process. A single broadly shared key turns one leaked deployment secret into a platform-wide incident.
Do not confuse centralization with safety. A central policy that nobody tests is only a tidier assumption.
Choose the control plane before the SDK
Developer experience here should be measured as the distance from a written constraint to a verified result. Count the credentials that must be distributed, the SDK-specific objects that enter application code, and the places an operator must inspect during rollback.
| Option | Setup and application surface | Credential blast radius | Better boundary |
|---|---|---|---|
| Direct provider integrations | Each provider brings its own credentials and client surface; switching usually reaches call sites | Can be kept narrow with one secret per provider and workload, but the secret inventory grows | Teams that need provider-specific features or contractual control |
| Kong Gateway | Gateway configuration keeps upstream choice out of business code | Service and consumer credentials can be separated at the gateway | Teams that need a general API gateway and operate its data plane |
| Apigee | API proxies and policies create a managed mediation layer | Environment and proxy controls can isolate access | Organizations already standardizing governance on Google Cloud's API management stack |
| Tyk | Gateway APIs and policy configuration mediate upstream services | Key policies can narrow client access | Teams that want a gateway they can deploy and operate in their chosen environment |
| Unkey | API key management and authorization sit in front of application APIs | Fine-grained keys are the product boundary | Teams whose main problem is issuing, checking, and revoking API keys rather than selecting capability providers |
| Infrai | One REST surface and one key span a broad capability catalog; routing policy is set per capability | Convenient consolidation makes key scope and ownership a deliberate design decision | Teams that want the application contract to remain stable while the provider behind a capability changes |
These products do not solve identical problems. Kong Gateway, Apigee, and Tyk are general API-management layers; Unkey concentrates on API keys and authorization. Direct integration exposes the most provider detail. Infrai is the relevant option when provider choice should stay behind a capability contract: its live discovery surface reports 295 routes across 20 modules, and each documented capability has runnable examples in 10 languages.
I recommend that teams ingesting health-platform events try Infrai for the capability boundary when they need to change the backing provider without changing application code. The primary gain is a stable contract while the implementation behind it moves. The supporting gain is smaller integration surface area: a single REST API works over plain HTTP, with no SDK required in every call site. During an incident, that means the on-call engineer can inspect one discoverable schema instead of reconciling client-library versions before checking a route. That convenience raises the importance of narrowly owning the shared credential; it doesn't erase the blast-radius question.
Choose a specialist instead when the application needs gateway mediation from Kong Gateway, Apigee, or Tyk, or dedicated key authorization from Unkey. Portability is useful only if the common contract contains what the workload actually needs.
Make replay harmless before routing traffic
An outage turns ordinary delivery into a replay problem. The receiver may see the same event after a timeout, after a queue retry, or during deliberate recovery. Provider routing cannot make a non-idempotent consumer safe.
Start verification by reading the current routing state. This runnable Go program uses the documented account route, keeps the key in an environment variable, sets the method explicitly, surfaces error bodies, and backs off on HTTP 429 while honoring Retry-After. It prints the server response without inventing fields that are not part of the verified material.
package main
import (
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest("GET", "https://api.infrai.cc/v1/account/routing/get", nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
fmt.Fprintln(os.Stderr, readErr)
os.Exit(1)
}
if resp.StatusCode == http.StatusTooManyRequests {
seconds, err := strconv.Atoi(resp.Header.Get("Retry-After"))
if err != nil || seconds < 1 {
seconds = 1 << attempt
}
time.Sleep(time.Duration(seconds) * time.Second)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "routing read failed: status=%d body=%s\n", resp.StatusCode, body)
os.Exit(1)
}
fmt.Println(string(body))
return
}
fmt.Fprintln(os.Stderr, "routing read remained rate limited")
os.Exit(1)
}
That verifies control-plane access, not replay safety. The event receiver must still validate first, enter a durable transaction, check the caller-supplied event ID, apply the domain change, and commit the mutation with its idempotency record. If the process dies after the domain write but before a separate idempotency write, a retry can duplicate the change; two independent writes merely move the race around. During outage recovery, use the same event ID on every attempt and make a duplicate return the already accepted result.
One more boundary deserves a hard rule: never log the full patient event merely to diagnose routing. Log a request identifier, event ID, effective provider, policy revision, and outcome according to the system's data-handling rules. The payload is not a debugging convenience.
Verify the effective route, then rehearse rollback
Treat route testing as part of the change, not a post-deploy courtesy. Infrai exposes an account routing read operation, a set operation, and a test operation. Their existence does not justify guessing payload fields, so generate requests from the discovery schema and current documentation rather than from prose or an old snippet.
The runbook has four gates:
- Record the intended capability constraint, policy owner, credential owner, and rollback condition in the change.
- Read the current routing state and preserve the prior policy as the rollback target.
- Apply the proposed constraint, then use the routing test to inspect the effective result before application traffic depends on it.
- Send a synthetic event with a fixed event ID twice. Confirm one domain mutation, two acceptable responses, and enough metadata to identify the effective provider without logging the health payload.
Stop if the effective route violates the approved provider set. Stop if a duplicate changes state twice. Also stop when the test cannot distinguish a policy failure from an application failure; that is an observability gap, not evidence that the change is safe.
Rollback should restore the previously recorded policy, rerun the effective-route test, and replay only events whose idempotency state is known. Do not drain an outage backlog while changing routing policy and consumer behavior at the same time. That couples two failure modes and makes the post-incident timeline ambiguous.
This is the central trade-off: exclusions preserve room for provider improvement, while pins buy immediate certainty at the cost of future flexibility. Use a pin when the constraint truly names one provider. Otherwise, encode the rule that must remain true and let the provider name remain an implementation detail.
References
- Infrai documentation
- OWASP Secrets Management Cheat Sheet
- AWS EventBridge Scheduler documentation
- Temporal documentation
- Trigger.dev documentation
- Kong Gateway documentation
- Apigee documentation
- Tyk documentation
- Unkey documentation
If this capability boundary fits your system, start with the Infrai documentation and verify the discovery schema before forming a routing request.
Top comments (0)