DEV Community

ApexZ69
ApexZ69

Posted on

Marketplace Backend Triage API: Node.js Maps OpenAI, Anthropic, Google With One Key

TL;DR: For marketplace support triage, put a small Node.js proxy between the application and the models. Keep one logical model name in product code, resolve it against a live catalog on the server, and reject any answer that does not match the ticket schema. Direct OpenAI, Anthropic Claude, and Google Gemini integrations give maximum provider-specific control; a unified runtime is the cleaner choice when one credential, one bill, and one operational view matter more than provider-specific features.

Start with the ownership trade-off. The table is the field guide; the implementation follows.

Pick Best fit What the backend owns Main boundary
OpenAI API directly The team has standardized on OpenAI One provider client, schema checks, retries, and usage telemetry Adding Claude or Gemini creates another integration path
Anthropic API directly Claude is the deliberate product choice Anthropic request mapping plus the same local validation and telemetry A later provider change reaches more application code unless an adapter already exists
Google Gemini API directly Gemini is the deliberate product choice Gemini request mapping plus local validation and telemetry Multi-provider routing remains the team's job
Unified OpenAI-compatible runtime Several model families must sit behind one server credential Logical-name policy, catalog refresh, validation, and request-level observability The common surface cannot expose every provider-specific control

How should a Node.js backend proxy map one API key?

Pick a direct API when the provider itself is part of the design. If prompts, safety behavior, or model-specific controls are tuned around OpenAI, then an extra compatibility layer buys little. The same rule applies to Anthropic Claude and Google Gemini. A thin internal adapter is still useful because ticket objects should not depend on a vendor response shape.

Pick a unified runtime when the provider is an implementation detail. Infrai is one example because one API key covers all capabilities through one REST API, with one wallet and one bill. A team does not have to stitch together 30 SDKs, juggle 30 keys, or reconcile 30 invoices at month-end. Its OpenAI-compatible surface covers multi-vendor model routing, while its model catalog reports currently available model IDs. The API is genuinely self-describing, and the public discovery surface requires no key; that matters when a deployment validates its model mapping before accepting traffic. For this triage path, consistent per-call cost, vendor, latency, cache, and request metadata also gives the service one place to build logs and alerts.

There is no magic in the proxy. It is a policy boundary. The browser asks for support-triage; the server resolves that name, sends the ticket, validates the reply, and records what happened. Product code never receives the runtime key.

Choose based on who should own provider variation. Direct integrations leave it with your team. A unified runtime moves transport and credential consolidation behind a common surface, while your team still owns output correctness.

Make correctness a local contract

A support ticket is not successful merely because a model returned fluent text. The useful result has four fields: a stable ticket identifier, one of exactly three queues, an integer urgency score from 1 through 5, and a reason capped at 240 characters. Anything else is an error to observe, not a creative interpretation to pass downstream. These limits are application policy rather than model capabilities, which is why they belong in code the team controls.

This is the diagram in words: marketplace form -> Node.js proxy -> logical model map -> available model -> chat completion -> local validator -> support queue. Beside that path, emit request ID, resolved model, provider, latency, estimated cost, validation outcome, and retry count. Those dimensions answer the first incident question: did the transport fail, or did a nominally successful response violate the contract?

Keep provider selection separate from parsing. It is tempting to let a fallback model return a slightly different shape and normalize it later. That creates branches precisely where production diagnosis needs a crisp invariant. One schema for every model is easier to alert on.

No exceptions.

Short is good here.

Use the runtime's token counting and cost estimation before dispatch when limits affect the decision. A large ticket can be warned, rejected, or assigned to a different approved logical class before a chat call is made. Batch APIs are a reasonable fit for an offline backlog, but interactive ticket triage should begin with standard chat completions.

Implement catalog-driven routing and bounded retries

The example below uses the OpenAI TypeScript client against an OpenAI-compatible base URL. It loads available IDs from the catalog, verifies that the deployment's logical mapping still points to a served model, retries HTTP 429 responses with Retry-After when present, and validates the JSON locally. Set INFRAI_API_KEY, UNIFIED_API_BASE_URL, and SUPPORT_TRIAGE_MODEL in the server environment. The base URL belongs in deployment configuration for this unlinked example; the model setting is an ID selected from the live catalog at deploy time.

import OpenAI from "openai";
import { randomUUID } from "node:crypto";

type Ticket = {
  id: string;
  subject: string;
  body: string;
};

type Triage = {
  ticketId: string;
  queue: "buyer" | "seller" | "trust";
  urgency: number;
  reason: string;
};

type ModelCatalog = {
  data: Array<{ id: string; available: boolean }>;
};

const apiKey = process.env.INFRAI_API_KEY;
const baseURL = process.env.UNIFIED_API_BASE_URL;
const configuredModel = process.env.SUPPORT_TRIAGE_MODEL;

if (!apiKey || !baseURL || !configuredModel) {
  throw new Error(
    "INFRAI_API_KEY, UNIFIED_API_BASE_URL, and SUPPORT_TRIAGE_MODEL are required",
  );
}

const client = new OpenAI({
  apiKey,
  baseURL,
  maxRetries: 0,
});

const sleep = (milliseconds: number) =>
  new Promise((resolve) => setTimeout(resolve, milliseconds));

async function loadModel(): Promise<string> {
  const response = await fetch(`${baseURL}/ai/models`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (!response.ok) {
    throw new Error(`Model catalog failed: ${response.status} ${await response.text()}`);
  }

  const catalog = (await response.json()) as ModelCatalog;
  const selected = catalog.data.find(
    (model) => model.id === configuredModel && model.available,
  );
  if (!selected) {
    throw new Error(`Configured model is unavailable: ${configuredModel}`);
  }
  return selected.id;
}

function retryDelay(error: OpenAI.APIError, attempt: number): number {
  const raw = error.headers?.get("retry-after");
  const seconds = raw === null || raw === undefined ? Number.NaN : Number(raw);
  return Number.isFinite(seconds) ? seconds * 1_000 : 500 * 2 ** attempt;
}

function parseTriage(raw: string, ticketId: string): Triage {
  const value = JSON.parse(raw) as Partial<Triage>;
  const queues = new Set(["buyer", "seller", "trust"]);
  if (
    value.ticketId !== ticketId ||
    typeof value.queue !== "string" ||
    !queues.has(value.queue) ||
    !Number.isInteger(value.urgency) ||
    value.urgency! < 1 ||
    value.urgency! > 5 ||
    typeof value.reason !== "string" ||
    value.reason.length < 1 ||
    value.reason.length > 240
  ) {
    throw new Error("Model response failed the triage contract");
  }
  return value as Triage;
}

async function triage(ticket: Ticket): Promise<Triage> {
  const model = await loadModel();
  const requestId = `ticket-${ticket.id}-${randomUUID()}`;

  for (let attempt = 0; attempt < 4; attempt += 1) {
    try {
      const completion = await client.chat.completions.create(
        {
          model,
          messages: [
            {
              role: "system",
              content:
                "Return only JSON with ticketId, queue, urgency, and reason. " +
                "queue must be buyer, seller, or trust; urgency must be 1 through 5.",
            },
            { role: "user", content: JSON.stringify(ticket) },
          ],
        },
        { headers: { "Idempotency-Key": requestId } },
      );
      const content = completion.choices[0]?.message.content;
      if (!content) throw new Error("Model returned no triage content");
      return parseTriage(content, ticket.id);
    } catch (error) {
      if (!(error instanceof OpenAI.APIError) || error.status !== 429 || attempt === 3) {
        throw error;
      }
      await sleep(retryDelay(error, attempt));
    }
  }
  throw new Error("Unreachable retry state");
}

const result = await triage({
  id: "MKT-1842",
  subject: "Payout is missing after a completed sale",
  body: "The order shows delivered, but the seller balance has not changed.",
});

process.stdout.write(`${JSON.stringify(result)}\n`);
Enter fullscreen mode Exit fullscreen mode

Four attempts is a ceiling, not a promise to keep hammering. A 429 gets an exponential delay unless the service provides Retry-After; authentication failures, invalid requests, and schema failures surface immediately. The idempotency key stays stable across attempts for the same operation. The explicit trade-off is a potentially slower final failure in exchange for absorbing short rate-limit windows; the cap prevents one ticket from occupying a worker indefinitely.

Refreshing the catalog on every ticket is deliberately easy to understand, but wasteful. In a service, load it at startup or deploy time as the facts of availability change, retain the last validated mapping, and fail readiness when a required logical model cannot be resolved. Do not silently choose an arbitrary replacement. That decision changes product behavior.

Fail closed.

Observe the boundary, not just the model call

The useful before/after is sharp. Before the proxy, each feature logs vendor-shaped fragments and a generic “AI failed” message. After it, every request has the same logical model, resolved model, ticket ID, request ID, attempt count, validation result, and terminal status. If the runtime returns per-call vendor, latency, cost, and cache metadata, attach those fields too.

Alert on outcomes. A rising 429 rate indicates capacity or quota pressure. A rising schema-rejection rate with a flat transport-error rate points toward prompt or model behavior. An unavailable configured ID is a deployment-readiness failure. These signals should not be collapsed into one error counter because they have different owners and fixes.

Never log the ticket body by default. Marketplace messages can contain names, addresses, order details, or payment context. Log the ticket ID and contract outcome; put content sampling behind an explicit privacy policy.

Know the limits before expanding the proxy

This pattern solves text routing. It does not make every AI function interchangeable. On the referenced unified runtime, ASR appears in the model directory with available=false; real-time voice session key status is pending and limited to the western region. There is no dedicated moderation endpoint, so text or image moderation needs a chat model with a JSON-schema fallback. Image upscaling is Lanc-only.

Those boundaries matter only if the ticket system grows into voice, moderation, or image workflows, but they should be recorded before anyone assumes “one API” means identical readiness everywhere. For the current job, keep the scope narrow: catalog-backed text models, one validated ticket contract, and observable retries.

The proxy earns its place when it makes a failure class visible. If the team uses one provider and needs its specialized controls, stay direct. If several model families must serve the same marketplace contract, a catalog-driven Node.js boundary keeps credentials, mapping, validation, and telemetry in one accountable place.

Sources

Top comments (0)