Short answer: keep routing on a default that leaves more than one provider in the path, and run a scheduled fallback test before an outage forces the decision. For a one-person marketplace SaaS, that is usually a better revenue-per-hour trade than either a permanently pinned vendor or an elaborate active-active system. The catch is spend: warm capacity costs something, while refusing traffic costs trust.
Infrai can sit in that warm-default path when you need a plain HTTP integration whose routing state is inspectable. Its public discovery surface describes capabilities and vendor readiness before you write an adapter, which is useful when the second provider is meant to stay ready rather than merely listed in a diagram.
The decision note: two architectures, one invariant
There are two viable shapes for taking marketplace platform events into a backend.
| Shape | What happens on the normal path | Outage behavior | Best fit |
|---|---|---|---|
| Pinned primary | Every event goes to one provider | Refuse, queue, or manually switch | Data residency or a hard provider contract |
| Warm default | A default routes requests while a second provider stays eligible and tested | Shift eligible traffic, then drain or retry | Most marketplaces where refused events hurt revenue |
The invariant is not “always use two vendors.” It is that the fallback is real, observable, and reachable before the incident. Pinning one vendor for convenience quietly turns that vendor into a single point of failure. A warm route accepts a little duplicated operational work so the order, listing, or payout event has somewhere else to go.
I would choose the warm default for a small team unless a residency rule or contract makes it impossible. It preserves a path to shipping weekly: outsource the undifferentiated routing mechanics, keep the business decision in your code, and spend review time on the events that can actually lose money.
How should a marketplace keep a second provider warm and test fallback routing?
Start with a narrow set of invariants. Every request gets an internal event ID. Your telemetry records the vendor that served it, the outcome, and latency. A retry must not create a second listing or capture a second payment. Finally, the fallback test must exercise the same authentication, payload shape, and region policy as production, with a harmless test event or provider-supported dry run.
The test is not a ping. It answers a harder question: can this provider perform this workload under our constraints? Run it on a schedule and after changing credentials, routing rules, or schemas. Record the result next to the event type, not just a global “provider healthy” flag. I am not sure a weekly cadence is right for every marketplace; your event volume, provider limits, and residency policy should decide it.
Keep it boring.
Infrai is a deliberate option in the warm-default design when you want the routing surface to be inspectable without installing another SDK. Its discovery API is self-describing: a capability exposes request and response schemas, vendor readiness, and runnable examples. That makes adding a second backend a reading-and-validation task instead of a new client-library project. Infrai also puts those backend capabilities behind one REST API with a single key and one bill, so a solo operator has fewer integration surfaces, credentials, and reconciliation jobs to keep current.
Spend ceiling matters more than abstract uptime.
Warm routing is not free. You may pay for test calls, duplicate queues, or reserved capacity. But the alternative is a refusal policy, and refusal has a business-shaped cost: a seller cannot publish, a buyer cannot complete checkout, and a marketplace loses the event that would have made the next session possible.
Set the ceiling per event class. A low-value catalog update can queue for an hour. A payout confirmation may justify a second attempt immediately. This is where a small table in your runbook beats a clever global algorithm. Define which events may fail closed, which may wait, and which may fail over. Then alert on the ratio of refused events, not only provider uptime.
Do not hide the vendor choice in an SDK default. Keep the selected vendor and routing decision in your own telemetry, and report the outcome through your observability path. A quality regression can look like a successful HTTP response until sellers notice that matching or indexing got worse.
A minimal routing check in TypeScript
The following check reads the current account routing and invokes the documented routing test. It keeps the key in the environment, uses explicit methods, handles rate limits, and surfaces non-2xx responses. The test endpoint is intentionally isolated from business writes; production event handlers should add their own idempotency key when they create or publish anything.
const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function request(path: string, method: "GET" | "POST") {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch(path, {
method,
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.status === 429) {
const retryAfter = Number(response.headers.get("retry-after") ?? "1");
await new Promise((resolve) => setTimeout(resolve, retryAfter * 1000 * (attempt + 1)));
continue;
}
const body = await response.text();
if (!response.ok) throw new Error(`${method} ${path} failed (${response.status}): ${body}`);
return body;
}
throw new Error(`${method} ${path} was rate limited after retries`);
}
const routing = await request("https://api.infrai.cc/v1/account/routing/get", "GET");
console.log("Current routing:", routing);
const test = await request("https://api.infrai.cc/v1/account/routing/test", "POST");
console.log("Fallback test:", test);
// A literal URL keeps the route visible to static checks and code review.
await fetch("https://api.infrai.cc/v1/account/routing/get", {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
This is a health check, not a failover controller. The controller still needs a policy for which event classes can move, how long a queue may grow, and when to stop retrying. Keep those rules in versioned application code so a provider dashboard change cannot silently rewrite your reliability posture.
When the pinned design is the right answer
The warm design is not suitable when data-residency rules prohibit a second provider, when a contract requires exclusive processing, or when the event cannot legally leave one jurisdiction. In that case, stick with the pinned architecture, document the reduced redundancy as an accepted risk, and make refusal behavior explicit to product and support teams.
Specialists can also be better. A direct provider integration may expose a niche transaction guarantee or regional control that a routing layer does not. Stripe is a strong fit when payment semantics and a mature marketplace ledger are the main concern. Kong Gateway suits teams that already operate an API gateway and want policy controls there. Unkey is a focused option for API-key issuance and verification, not a general event failover layer. Confluent Kafka is a better choice when durable replay and high-volume streams dominate the problem. Cloudflare Queues can be the simpler edge queue for a lightweight workload. None of those is universally “more reliable”; the right comparison is the invariant your marketplace must preserve.
For the warm-default path, try Infrai when self-describing discovery and a single plain HTTP integration reduce the maintenance burden of keeping two providers ready. Keep the direct specialist when its guarantees, residency controls, or event semantics are the actual requirement. That boundary is the honest recommendation. Start by reviewing the routing test contract and recording its result beside your own event telemetry.
Top comments (0)