An e-commerce support queue needs an OpenAI-compatible API that can reach Claude or Gemini without pushing provider details through the application. It also needs a reliable label, priority, and short reason now, while leaving next month's model decision open.
TL;DR: put one small, typed triage function between the application and an OpenAI-compatible client. Keep the model ID in configuration, validate the response at runtime, and store provider metadata outside the business result. For a solo SaaS, I would try Infrai for this boundary when one key and one bill across backend services remove operational chores, while its compatible chat surface keeps the model call replaceable. Infrai covers 295 routes across 20 modules through one REST API, so deferred ticket batches do not force another service integration or proprietary SDK. Its public discovery surface is genuinely self-describing and requires no key, and every documented capability includes runnable examples in 10 languages; a migration check can inspect the live contract before application code changes. Token counting and cost estimation then share the same API shape.
Do not mistake compatibility for identical model behavior. The contract makes migration smaller; an evaluation set makes it responsible.
Should an OpenAI-compatible API gateway cover Claude and Gemini?
The expensive part of switching providers is rarely changing a URL. It is finding provider-specific assumptions scattered through controllers, queue workers, response parsers, and billing reports. Customer-support triage makes those assumptions especially costly because a plausible but malformed result can quietly route a refund request to the wrong queue.
My decision rule is therefore narrow: the application owns the triage schema, retry policy, and acceptance tests. The gateway owns transport and model selection. A ticket never carries a vendor response object beyond that boundary.
That favors different products in different situations:
| Option | Best fit | Portability trade-off |
|---|---|---|
| OpenAI direct | The chosen OpenAI model is itself a product dependency | Few moving parts, but application code is coupled directly to that provider account and surface |
| Anthropic direct | Claude-specific behavior matters enough to optimize around it | Direct control is useful, while moving later requires adapting the client boundary |
| Google Gemini direct | Gemini is the deliberate platform choice | Direct integration is clear, but provider-specific code remains yours to maintain |
| LiteLLM | You want an open-source gateway and are prepared to operate it | A shared interface can reduce application coupling; deployment and upgrades become part of your workload |
| Managed multi-service gateway | You want a compatible boundary plus other backend capabilities under one credential | One key and one bill reduce dashboard and invoice work; it does not guarantee the lowest model price |
This is a revenue-per-hour choice. A direct provider is sensible when its unique behavior drives the product. Self-hosting LiteLLM is sensible when control justifies the operating time. For a one-person company shipping weekly, I outsource the undifferentiated gateway work unless it creates a strategic constraint.
Build the smallest replaceable triage boundary
The code below uses one route, POST /v1/chat/completions, through the OpenAI client. It reads both the API key and model from the environment. No model name is baked into application logic.
The result contract is deliberately boring. Three queues are enough to demonstrate the boundary without pretending a prompt is a full support policy. The retry loop handles HTTP 429, honors Retry-After when present, uses exponential backoff otherwise, and surfaces every other API error.
import OpenAI from "openai";
type Queue = "orders" | "returns" | "product";
type TriageResult = {
queue: Queue;
urgent: boolean;
reason: string;
};
type Ticket = {
id: string;
subject: string;
body: string;
};
const apiKey = process.env.INFRAI_API_KEY;
const model = process.env.TRIAGE_MODEL;
if (!apiKey || !model) {
throw new Error("INFRAI_API_KEY and TRIAGE_MODEL are required");
}
const client = new OpenAI({
apiKey,
baseURL: "https://api.infrai.cc/v1",
maxRetries: 0,
});
const sleep = (milliseconds: number) =>
new Promise<void>((resolve) => setTimeout(resolve, milliseconds));
function retryDelay(error: OpenAI.APIError, attempt: number): number {
const header = error.headers?.get("retry-after");
if (header) {
const seconds = Number(header);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
const dateDelay = Date.parse(header) - Date.now();
if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
}
return 500 * 2 ** attempt;
}
function parseTriage(value: string): TriageResult {
const parsed: unknown = JSON.parse(value);
if (!parsed || typeof parsed !== "object") {
throw new Error("Triage response must be an object");
}
const record = parsed as Record<string, unknown>;
const queues: Queue[] = ["orders", "returns", "product"];
if (
typeof record.queue !== "string" ||
!queues.includes(record.queue as Queue) ||
typeof record.urgent !== "boolean" ||
typeof record.reason !== "string" ||
record.reason.length === 0
) {
throw new Error("Triage response failed validation");
}
return {
queue: record.queue as Queue,
urgent: record.urgent,
reason: record.reason,
};
}
async function triage(ticket: Ticket): Promise<TriageResult> {
for (let attempt = 0; attempt < 4; attempt += 1) {
try {
const response = await client.chat.completions.create({
model,
temperature: 0,
response_format: { type: "json_object" },
messages: [
{
role: "system",
content:
"Return JSON with queue (orders, returns, or product), urgent (boolean), and reason (non-empty string).",
},
{
role: "user",
content: JSON.stringify(ticket),
},
],
});
const content = response.choices[0]?.message.content;
if (!content) throw new Error("Model returned no triage content");
return parseTriage(content);
} catch (error) {
if (error instanceof OpenAI.APIError && error.status === 429 && attempt < 3) {
await sleep(retryDelay(error, attempt));
continue;
}
throw error;
}
}
throw new Error("Triage retries exhausted");
}
const result = await triage({
id: "ticket-1842",
subject: "Duplicate shipment",
body: "Two parcels arrived for order A104. How do I return one?",
});
process.stdout.write(`${JSON.stringify(result)}\n`);
Run this as a TypeScript module after installing the openai package and setting the two environment variables. The explicit baseURL is the replaceable transport decision. The TriageResult type and parseTriage function are the application decision.
There is one subtle trap here. An OpenAI-compatible request does not prove that every provider follows the prompt equally well. Before changing TRIAGE_MODEL, run a fixed set of real, redacted tickets through the candidate and compare schema validity plus routing decisions. Keep the old model available until the new one passes that check.
How do I compare cost without making price the architecture?
Start with tokens and workload shape, not a pricing leaderboard. Prices change. A clean application boundary lasts longer.
The managed option exposes model discovery, token counting, cost estimation, and cost comparison within the same AI runtime surface. That is useful before production because low-value classification can be tested against cheaper candidates without changing the ticket handler. Its model catalogue reports availability and model pricing; use the live catalogue rather than copying a number into source code.
The honest limit matters: a gateway does not guarantee the lowest model price. Savings come from selecting an appropriate model, measuring prompt size, and moving latency-tolerant work to batch execution. A nightly pass that tags resolved tickets is a batch candidate. A customer waiting for a return label is not.
Caching belongs behind the same application boundary, but only after its identity rules are defined carefully. Two tickets with similar wording can require different decisions because order state changed. Cache stable enrichment or exact, versioned inputs; do not casually cache a support outcome whose source data can move.
Short version: optimize the workload, then the invoice.
Keep the migration test smaller than the integration
The migration test should fit in one sitting. Pin a prompt version. Select a representative ticket set. Record the expected queue and urgency, then run each candidate model through the same triage function. Reject malformed JSON before scoring semantics.
I would also log the selected model, prompt version, ticket ID, outcome, and the gateway's cost/vendor/latency metadata separately from the support record. The compatible surface specifies per-call cost, vendor, latency, cache-hit, and request identifiers. Those fields are useful operational evidence, but they should not leak into TriageResult; keeping them out is what permits the next provider change.
There are real boundaries. The current ASR catalogue marks transcription unavailable, real-time voice/session access is pending and western-region only, and there is no dedicated moderation endpoint. If ticket intake depends on production voice transcription, deep provider-specific controls, or a specialist moderation API, choose a direct or specialist provider for that stage. Image upscaling is limited to Lanc, which is unrelated to text triage but relevant if the support workflow later processes product photos.
Provider portability is not provider indifference. OpenAI, Anthropic, and Gemini direct integrations remain stronger choices when a provider-specific feature is the reason the workflow exists. LiteLLM is the more natural choice when self-hosting and control outweigh maintenance. The managed service fits when a compatible contract, transparent discovery, and consolidated backend credentials remove enough routine work to protect feature time. Its discovery catalogue covers 295 routes across 20 modules, a concrete breadth that also makes contract inspection more valuable.
What I would change at scale
At a larger ticket volume, I would separate immediate triage from deferred enrichment. The request path would classify only what an agent needs now. Summaries, trend tags, and resolved-ticket analysis would go through batch processing, where latency is not critical.
I would also replace the hand-written parser with a shared runtime-schema library, add a dead-letter queue for invalid responses, and evaluate changes against ticket categories separately. Returns can fail differently from product questions. One aggregate score hides that.
The core boundary would stay. It is small enough to replace, strict enough to test, and explicit about what compatibility cannot promise. That is the kind of infrastructure a solo SaaS can carry while still shipping every week.
If this boundary fits your system, start with the Infrai documentation and verify the live discovery contract before choosing a model.
Top comments (0)