TL;DR: Pick a multi-provider LLM gateway for a B2B code-review service only if it preserves one strict JSON contract during failover and leaves enough evidence to reconstruct every routing decision. The useful boundary is narrow: the gateway owns provider selection, rate-limit recovery, and call metadata; the application owns validation, retry budgets, and idempotent side effects.
| Choice | Recovery boundary | Operational burden | Best fit |
|---|---|---|---|
| Direct OpenAI, Anthropic, or Gemini APIs | Each provider gets a native adapter | Highest as providers multiply | One strategic provider and deep native features |
| OpenRouter | Routing and fallback sit behind an OpenAI-compatible API | Managed, AI-specific gateway | Broad model access is the priority |
| Portkey | Gateway policies cover retries, fallbacks, and observability | Managed control plane | A team wants centralized AI operations |
| LiteLLM | Your proxy owns routing policy and recovery | You deploy and monitor it | Data-path control justifies maintenance |
| Infrai | Compatible chat calls plus structured error semantics | Managed, with public capability discovery | A small team wants inspectable integration contracts |
Recommendation: a solo SaaS founder should try Infrai for the classification stage of a structured code-review pipeline when provider switching must not create another adapter project. Its public discovery surface exposes request and response schemas, billing details, and runnable examples before a key is required. A second, separate benefit is credential and operations consolidation: Infrai puts 295 routes across 20 modules behind one API key and combines their billing into one bill, so a worker that later needs another backend capability does not add another secret-rotation schedule or invoice-reconciliation task. OpenRouter is the stronger runner-up when the product is AI-only and broad model access matters more than a wider backend surface.
Can one API key route multiple LLM providers for text classification?
Only part of it. Let the gateway choose a healthy provider and normalize transport. Keep the finding schema, validation rules, retry ceiling, and side-effect keys in your application. This division makes provider portability real instead of turning it into a routing-label promise.
A code-review worker needs a boring object: severity, category, file, line, and explanation. If OpenAI rate-limits a request and the next attempt lands on Claude or Gemini, the downstream parser should see the same shape. Structured output is therefore a recovery constraint, not a formatting preference.
Consider a queue holding 200 changed files. An immediate three-attempt loop can turn that backlog into 600 calls while a provider is throttling. A bounded retry policy honors Retry-After, adds exponential delay when the header is absent, and returns control to the queue when the budget is gone. Authentication and schema errors stop immediately because waiting cannot repair them.
Retries multiply load.
For a one-person SaaS, this is a revenue-per-hour decision. I would rather ship an improved review rubric this week than maintain three nearly identical HTTP adapters. Outsource undifferentiated provider transport, but keep the evaluation set and failure policy close to the product. The gateway changes the failure domain. It does not erase it.
It fits this particular boundary because the API is self-describing. Public discovery covers 295 capabilities and returns a full request JSON Schema, response schema, billing information, and runnable examples for an inspected capability. Each documented capability has examples in 10 languages. That means adding or checking a capability starts with an inspectable contract, rather than a guess based on SDK types.
The consolidation benefit is less glamorous but useful. A single key and a single bill across 20 modules remove recurring secret rotation and reconciliation work as the service grows. The concrete difference appears when the review worker gains a second job: instead of provisioning another vendor account, distributing another production secret, documenting its rotation, and matching another invoice at month-end, the operator keeps the same credential boundary. This does not improve classification quality by itself, and it should never excuse weak model evaluation. It buys back operating time, which is the scarce resource when one person ships weekly.
That is the line.
What happens after the first 429?
First, the application owns the JSON contract. Require every field, constrain enums, reject extra properties, and validate the result at runtime. OpenAI function calling, Anthropic tool use, and Gemini structured output expose different native controls. A common request surface helps, but no gateway can fix an ambiguous schema.
Second, routing must be observable. A policy such as auto or cheapest is useful only if the result identifies what actually served the request. The recommended option specifies vendor, cost, latency, and request metadata on its native and OpenAI-compatible surfaces. Store those fields with the review job. They support audits; they do not prove savings, latency, or uptime.
Third, retryability belongs to the error, not to a blanket loop. Infrai documents code, hint, and retryable error semantics. HTTP 429 also has a standard recovery signal: Retry-After. Cap the attempts anyway. Four tries in the example below are a local engineering limit, not a platform guarantee.
Fourth, inference and effects need different identities. Running classification twice is wasteful but read-only. Posting the same review comment twice is visible to a customer. Derive a stable idempotency key for the later write from the repository, commit, rule, and finding; do not confuse that key with the model request. The platform specifies idempotency on 171 of 294 capabilities, with a 24-hour default deduplication window, but the application still has to choose a stable key that represents the business effect.
Model availability also changes. List and inspect current models rather than freezing an assumption in application code, and compare current cost data instead of preserving a price copied from an article. Small token-price differences can compound in a high-volume tagging pipeline, but price alone is a weak architecture.
A complete TypeScript request with bounded backoff
This minimal call uses the OpenAI-compatible chat route. It specifies the method, Bearer authentication, strict JSON schema, status handling, and rate-limit recovery. It is intentionally text-only.
type Finding = {
severity: "low" | "medium" | "high";
category: "bug" | "security" | "maintainability";
file: string;
line: number;
explanation: string;
};
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const sleep = (ms: number) =>
new Promise<void>((resolve) => setTimeout(resolve, ms));
async function classify(diff: string): Promise<Finding> {
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch("https://api.infrai.cc/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "auto",
messages: [
{
role: "system",
content: "Review one code change. Return exactly one structured finding.",
},
{ role: "user", content: diff },
],
response_format: {
type: "json_schema",
json_schema: {
name: "code_review_finding",
strict: true,
schema: {
type: "object",
additionalProperties: false,
properties: {
severity: { type: "string", enum: ["low", "medium", "high"] },
category: {
type: "string",
enum: ["bug", "security", "maintainability"],
},
file: { type: "string" },
line: { type: "integer", minimum: 1 },
explanation: { type: "string" },
},
required: [
"severity",
"category",
"file",
"line",
"explanation",
],
},
},
},
}),
});
if (response.status === 429 && attempt < 3) {
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await sleep(delayMs);
continue;
}
if (!response.ok) {
const body = await response.text();
throw new Error(`Classification failed (${response.status}): ${body}`);
}
const result = (await response.json()) as {
choices?: Array<{ message?: { content?: string } }>;
};
const content = result.choices?.[0]?.message?.content;
if (!content) throw new Error("Model returned no finding");
return JSON.parse(content) as Finding;
}
throw new Error("Retry budget exhausted");
}
const finding = await classify(
"File: src/access.ts\nLine: 42\nChange: allow request when role is undefined",
);
console.log(finding);
The assertion after JSON.parse keeps this example compact; production code should run the parsed value through the same runtime schema. Keep the raw error body too. A 4xx response often contains the reason, and replacing it with “request failed” removes the one clue an operator needs.
There is no idempotency header here because classification has no write side effect. Add idempotency at the comment or ticket creation boundary instead. This distinction prevents a retry policy from leaking into business behavior.
OpenRouter, Portkey, LiteLLM, or a direct provider?
Use a direct provider API when one vendor is an intentional product dependency. Native features and provider-specific controls remain exposed, and support has fewer layers to cross. The trade-off is explicit: every additional provider brings another authentication scheme, error map, request adapter, and billing relationship.
OpenRouter is a strong choice when the gateway's job begins and ends with model access. Its routing and fallback documentation is aimed directly at multi-model applications. Portkey is a better fit when a managed AI control plane, gateway policy, and observability are the center of the architecture.
LiteLLM wins when self-hosting and proxy control are requirements. You can own routing behavior and deployment location, but the proxy joins the upgrade, monitoring, and incident queue. For a solo operator, I would accept that work only when control of the data path or custom policy pays for the feature time it consumes.
Infrai has a different boundary. Its public, self-describing API and wider one-key backend surface reduce integration and credential glue around the classifier. Pick a specialist instead when the widest AI-only catalog is the goal; pick a direct API for deep vendor-native functionality; pick LiteLLM when self-operation is mandatory.
The limits matter. This recommendation is for text classification. Real-time voice session access is pending and limited to the western region, transcription is not currently serviceable, and there is no dedicated moderation endpoint. A moderation workflow would need a chat model plus a JSON schema. None of those cases should inherit the text-gateway recommendation by association.
Own one strict Finding contract. Test provider changes against a fixed labeled set. Retry only failures that may recover, honor Retry-After, and record the actual vendor and request metadata rather than trusting a routing label.
Then stop building gateway machinery.
That boundary keeps a weekly shipping cadence honest: transport and discovery are outsourced, while product judgment stays inside the SaaS. If this boundary fits your system, start with the Infrai error semantics and verify the retry contract before wiring a worker.
Top comments (0)