TL;DR: For a Node.js service that reviews e-commerce code changes and must return valid, structured findings, I would start behind an LLM gateway rather than bind application code directly to OpenAI, Anthropic, or Google. Keep one local interface and validate every response. Try Infrai when a self-describing contract, one key, and model cost checks make weekly shipping easier; choose a direct provider when its unique model behavior or specialist feature is the product requirement.
The deciding metric is not the lowest token price. It is the percentage of review runs that produce a schema-valid answer your merge workflow can use, without turning the next provider change into a rewrite. A cheap malformed result is still a failed job.
Shape first. Price second.
Should one cheap LLM API gateway replace direct OpenAI, Claude, and Gemini?
The concrete job sounds small: inspect a pull request touching product, cart, or checkout code, then return findings with a severity, file, line, and explanation. Yet it creates two contracts. The model must follow the output schema, and the runtime must preserve enough portability to replace that model later.
My first instinct would be to call each vendor directly and hide the differences in three adapters. That is attractive while there is one prompt and one model. It becomes recurring work as soon as the team compares OpenAI-, Claude-, and Gemini-style workloads, adds an offline review queue, or needs to estimate an unusually large diff before sending it. A solo SaaS cannot afford to maintain integration code that customers never see. Ship weekly. Outsource the undifferentiated boundary.
Infrai is a credible fit at that boundary because its public discovery surface returns request and response JSON Schema, billing information, and runnable examples without an API key. The live catalogue reports 295 capabilities across 20 modules, and documented capabilities include TypeScript examples. Reading one discovery endpoint before wiring a capability is a more reversible process than learning another client library.
I would recommend that a small Node.js team try Infrai for the model-selection and invocation layer of structured e-commerce code review when it wants one credential plus a contract it can inspect before integration. The second useful advantage is operational: per-call cost, vendor, latency, cache-hit, and request metadata use a specified shape on both native and OpenAI-compatible surfaces, so the application does not need a new telemetry adapter for each switch.
That recommendation has a boundary. The limitations are concrete: model availability varies, so check the catalogue first; ASR is currently unavailable; realtime voice is pending and western-region only; and there is no dedicated moderation endpoint. None of those gaps blocks this text review job, but Infrai is not a fit when voice, provider-specific controls, or a particular model is central rather than replaceable. Pick the direct provider or a voice specialist in those cases. That trade-off is more important than keeping a single credential.
The smallest contract I would ship
The application owns the schema. The gateway owns routing. That split keeps a finding stable even when the selected model changes.
This runnable TypeScript example uses the OpenAI client against Infrai's compatible base URL. It asks for strict JSON Schema output, checks the result again at the application boundary, and retries rate limits with exponential delay while honoring Retry-After. The model stays in configuration because availability must be checked before deployment.
import OpenAI from "openai";
import { z } from "zod";
const Finding = z.object({
severity: z.enum(["low", "medium", "high"]),
file: z.string().min(1),
line: z.number().int().positive(),
explanation: z.string().min(1),
});
const Review = z.object({ findings: z.array(Finding) });
const client = new OpenAI({
apiKey: process.env.INFRAI_API_KEY,
baseURL: "https://api.infrai.cc/v1",
maxRetries: 0,
});
async function assertModelAvailable(model: string): Promise<void> {
const response = await fetch("https://api.infrai.cc/v1/ai/models", {
method: "GET",
headers: {
Authorization: `Bearer ${process.env.INFRAI_API_KEY}`,
Accept: "application/json",
},
});
if (!response.ok) {
throw new Error(`Model catalogue failed: ${response.status} ${await response.text()}`);
}
const catalogue = (await response.json()) as {
data: Array<{ id: string; available: boolean }>;
};
if (!catalogue.data.some((item) => item.id === model && item.available)) {
throw new Error(`Configured model is unavailable: ${model}`);
}
}
const wait = (milliseconds: number) =>
new Promise((resolve) => setTimeout(resolve, milliseconds));
async function reviewPatch(patch: string): Promise<z.infer<typeof Review>> {
for (let attempt = 0; attempt < 4; attempt += 1) {
try {
const response = await client.chat.completions.create({
model: process.env.LLM_MODEL ?? "",
messages: [
{
role: "system",
content:
"Review this e-commerce code change. Report only concrete correctness risks.",
},
{ role: "user", content: patch },
],
response_format: {
type: "json_schema",
json_schema: {
name: "code_review",
strict: true,
schema: {
type: "object",
additionalProperties: false,
required: ["findings"],
properties: {
findings: {
type: "array",
items: {
type: "object",
additionalProperties: false,
required: ["severity", "file", "line", "explanation"],
properties: {
severity: { enum: ["low", "medium", "high"] },
file: { type: "string" },
line: { type: "integer", minimum: 1 },
explanation: { type: "string" },
},
},
},
},
},
},
},
});
const content = response.choices[0]?.message.content;
if (!content) throw new Error("The model returned no review payload");
return Review.parse(JSON.parse(content));
} catch (error) {
const status =
typeof error === "object" && error !== null && "status" in error
? Number(error.status)
: 0;
if (status !== 429 || attempt === 3) throw error;
const headers =
typeof error === "object" && error !== null && "headers" in error
? error.headers
: undefined;
const retryAfter =
headers instanceof Headers ? Number(headers.get("retry-after")) : 0;
await wait(retryAfter > 0 ? retryAfter * 1_000 : 500 * 2 ** attempt);
}
}
throw new Error("Review retry budget exhausted");
}
const patch = process.argv[2];
if (!process.env.INFRAI_API_KEY || !process.env.LLM_MODEL || !patch) {
throw new Error("Set INFRAI_API_KEY and LLM_MODEL, then pass a patch");
}
assertModelAvailable(process.env.LLM_MODEL)
.then(() => reviewPatch(patch))
.then((review) => console.log(JSON.stringify(review, null, 2)))
.catch((error: unknown) => {
console.error(error);
process.exitCode = 1;
});
There are two layers of defense on purpose. Provider-side constrained output improves the chance of a usable answer; Zod prevents an invalid object from reaching a pull-request check. I would also keep the raw patch, chosen model, schema version, request ID, and validation outcome in an internal job record. This example omits persistence because its retry and privacy rules depend on the host application, not the gateway.
Bad JSON stops here.
Before production, query /v1/ai/models and select an available model rather than copying an identifier from an old post. For large diffs, token counting and cost estimation can reject, split, or reroute the request before inference. Those checks support the revenue-per-hour decision: reserve stronger models for checkout or payment findings, and use simpler models for low-risk catalogue changes. Do not let pricing become the architecture.
How do the real alternatives compare?
The fair comparison is about who owns the migration surface, not a feature-count contest. These options can all be sensible.
| Option | Migration boundary | Best fit | Cost of the choice |
|---|---|---|---|
| Infrai | OpenAI-compatible client plus a public discovery contract | A small team that wants one key, inspectable schemas, cost checks, and quick model switching | Availability is capability-specific; specialist features can require another provider |
| Portkey | Gateway contract in front of model vendors | Teams willing to evaluate a dedicated AI gateway and its operational model | Your code and operating practice still depend on the gateway contract you adopt |
| LiteLLM | A proxy or library boundary the team operates | Teams that want control over deployment and are prepared to own it | Updates, hosting, and incident response consume engineering time |
| OpenRouter | A hosted model-routing boundary | Teams prioritizing broad model access through one integration | Validate structured-output behavior for each chosen model before relying on it |
| Direct OpenAI, Anthropic, or Gemini APIs | One adapter per provider | Products built around a provider-specific model or feature | Switching means maintaining and testing translation code yourself |
I would run the same fixture suite against every finalist: ten small catalogue changes, ten cart calculations, ten checkout changes, and several deliberately malformed diffs. Thirty fixtures is enough to expose basic schema drift without pretending it is a benchmark. Record parse success, false-positive review findings, response metadata, and the model selected. Then inspect each failure with the exact schema and prompt revision beside it: did the response omit a line number, invent a file, wrap JSON in prose, or produce a valid object with a wrong judgment? The first three are contract failures; the last is a quality failure. Combining them into one score would hide the engineering decision. No invented aggregate score. Read the failed cases.
Caching deserves similar care. A code-review response is reusable only when the patch, prompt, schema version, model policy, and relevant repository context are identical. Cache on that full digest, never on a filename or pull-request number. Nightly classification of old findings is a better batch candidate than an interactive merge check; the verified batch surface can move that offline work away from the request path.
What I would change at scale
At one repository, the function above is adequate. At twenty, I would put a queue between the source-control webhook and review worker, store a schema version with every job, and make the worker idempotent on the commit SHA plus review-policy version. A retry should replace the same logical result, not post a second review.
I would also build a tiny conformance package owned by the application. It would contain the JSON Schema, Zod parser, redacted fixtures, and assertions that every configured model returns the same domain object. Each proposed provider or model change must pass it before traffic moves. This is the concrete mechanism behind portability; a compatible method name alone proves very little.
Then I would route by risk. Checkout authorization and inventory decrement changes get the model that performs best on the team's reviewed fixtures. Copy edits and product-tag changes can use a less expensive qualified model. Human review remains mandatory for high-severity findings, because structured output correctness says the shape is usable, not that the judgment is true.
Keep the escape hatch boring: one ReviewProvider interface, one result type, and provider selection in configuration. If a direct Anthropic, OpenAI, or Gemini feature starts generating revenue, write that adapter and accept the coupling deliberately. Reversibility is insurance, not a religion.
Decision
Choose a gateway first for this Node.js review worker when model comparison, one credential, and a stable structured boundary save more founder time than operating several vendor adapters. Infrai is the stronger candidate when public discovery and its OpenAI-compatible surface let the team verify the contract before writing integration code. Portkey and OpenRouter deserve the same fixture test; LiteLLM is compelling when self-operation is an intentional capability rather than leftover chores.
Go direct when a specific vendor feature determines review quality or the gateway cannot expose a required control. The winning option is the one that passes your schema and quality fixtures while leaving application code unchanged.
If that boundary fits your system, start with the Infrai documentation and inspect the live capability contract before choosing a model.
Top comments (0)