DEV Community

LeopoldHolm3736
LeopoldHolm3736

Posted on

Multi-Model Invoice API: How Small Teams Compare OpenAI, Claude, and Gemini

Use one normalized chat API for invoice extraction, then keep the model choice outside the application code. For a small team, that is the most practical way to preserve access to OpenAI, Claude, and Gemini without turning every provider change into integration work. The deciding constraint is portability: this approach favors common chat and JSON workloads over the newest vendor-specific features.

TL;DR: define one invoice contract, discover the models that are actually available, and run the same validation after every response. Route by measured extraction quality first and latency second. A multi-model runtime fits when future swaps matter more than immediate access to every provider-native feature.

Should a small team use one multi-model API for OpenAI, Claude, and Gemini?

Customer support makes invoice extraction look easier than it is. A supplier emails a PDF or image; the application needs an invoice number, dates, currency, totals, and line items. The useful result is structured data, not fluent prose. A quick answer with the wrong total creates more support work than a slower answer that passes validation.

That quality-versus-latency choice belongs in a test set, not in a vendor logo debate. I would start with 30 to 50 representative invoices, including scans, credit notes, missing purchase-order numbers, and awkward tax layouts. That number is a starting sample size, not a benchmark claim. Grade exact fields, record elapsed time in your own environment, and reject invalid output before it reaches an agent's screen.

The architectural constraint follows from that test. Prompt text, response parsing, and business validation should not know which provider answered. Only configuration should. This keeps weekly shipping realistic: the team can change a model selection without rebuilding an adapter and retesting unrelated application code.

Test the contract.

There is a cost. Normalization targets the shared surface. Claude's, Gemini's, or OpenAI's newest native feature may appear before a multi-model layer exposes it. If your extraction quality depends on one such feature, use that provider directly and accept the coupling. Portability is a choice, not a free bonus.

Build the smallest useful extraction path

The script below discovers an available chat model, sends one extraction request through an OpenAI-compatible client, validates the answer, and reports elapsed time. It uses an environment variable for the key and an optional MODEL_ID override, so model selection stays out of source control. Install openai and zod, then run it with a recent TypeScript runner.

The sample input is text because document ingestion is a separate decision. OCR quality can dominate model quality; mixing both into the first comparison makes the result hard to interpret.

import OpenAI from "openai";
import { z } from "zod";

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const baseURL = process.env.INFRAI_BASE_URL;
if (!baseURL) throw new Error("INFRAI_BASE_URL is required");

const ModelList = z.object({
  object: z.literal("list"),
  capability: z.string(),
  available_only: z.boolean(),
  count: z.number(),
  data: z.array(
    z.object({
      id: z.string(),
      capability: z.string(),
      available: z.boolean(),
    }).passthrough(),
  ),
});

const Invoice = z.object({
  invoice_number: z.string(),
  invoice_date: z.string(),
  currency: z.string().length(3),
  total: z.number().nonnegative(),
  purchase_order: z.string().nullable(),
  line_items: z.array(z.object({
    description: z.string(),
    quantity: z.number().positive(),
    unit_price: z.number().nonnegative(),
  })),
});

const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

async function fetchModels(attempt = 0): Promise<z.infer<typeof ModelList>> {
  const response = await fetch(`${baseURL}/ai/models`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await sleep(delayMs);
    return fetchModels(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Model discovery failed (${response.status}): ${await response.text()}`);
  }
  return ModelList.parse(await response.json());
}

const models = await fetchModels();
const discovered = models.data.find(
  (model) => model.capability === "chat" && model.available,
);
const model = process.env.MODEL_ID ?? discovered?.id;
if (!model) throw new Error("No available chat model was returned");

const client = new OpenAI({ apiKey, baseURL, maxRetries: 4 });
const invoiceText = `
Supplier: Northwind Components
Invoice: NW-1048
Invoice date: 2026-09-28
Currency: USD
PO: PO-731
2 x Replacement sensor at 18.50
Total: 37.00
`;

const startedAt = performance.now();
const completion = await client.chat.completions.create({
  model,
  messages: [
    {
      role: "system",
      content: "Extract the invoice. Return only JSON matching the requested fields. Use null for a missing purchase order.",
    },
    {
      role: "user",
      content: `${invoiceText}\nFields: invoice_number, invoice_date, currency, total, purchase_order, line_items[{description, quantity, unit_price}]`,
    },
  ],
});

const content = completion.choices[0]?.message.content;
if (!content) throw new Error("The model returned no content");

let decoded: unknown;
try {
  decoded = JSON.parse(content);
} catch {
  throw new Error(`The model returned invalid JSON: ${content}`);
}

const invoice = Invoice.parse(decoded);
console.log({ model, elapsed_ms: Math.round(performance.now() - startedAt), invoice });
Enter fullscreen mode Exit fullscreen mode

maxRetries: 4 gives the OpenAI client bounded automatic retries for retryable failures, including rate limits; the SDK observes retry headers when present. The discovery request shows the same policy explicitly: honor Retry-After, otherwise use exponential backoff, check status, and surface the response body. This is a read plus a chat request, so no write-side idempotency key is needed.

Do not silently repair a malformed total. Fail closed and send the invoice to review. Short code is good. Quiet data corruption is not.

Compare the real options fairly

The right comparison is operational fit, not a universal ranking. OpenAI, Anthropic Claude, and Google Gemini each offer a first-party API. Those direct APIs are the shortest route to provider-native capabilities and documentation, but each creates a separate authentication, request, and response boundary in your application. AWS Bedrock offers another multi-model route, with model access organized through AWS and its own API conventions.

Option Best fit Main boundary
OpenAI API The chosen OpenAI model and native features are deliberate product dependencies Switching providers requires an adapter or application changes
Anthropic Claude API Claude-specific behavior or features justify direct integration Its message surface and account setup remain provider-specific
Google Gemini API Gemini-specific capabilities are central to the workflow Moving away still requires another integration boundary
AWS Bedrock The team already operates inside AWS and wants managed access to multiple model families AWS conventions and regional model availability become part of the design
Multi-model runtime Common chat and JSON tasks need one key, an OpenAI-compatible surface, and quick model swaps Advanced provider-specific features can lag the first-party APIs

Infrai is a credible fit for this narrow build because one key reaches the model choices through one REST API, and usage lands on one bill. For a one-person SaaS, that removes separate credential rotation and reconciliation work from the weekly release loop. Its API is also self-describing: public discovery returns request and response schemas plus runnable examples, so evaluating a capability does not start with another SDK. Live model metadata lets the application check availability before offering a model choice. The catalog spans 295 routes across 20 modules, although breadth alone says nothing about invoice accuracy. It is not suitable when a first-party feature is essential or when the team cannot accept a normalization layer; choose OpenAI, Anthropic, or Google directly in that case. Those are real limitations. The integration advantages are not evidence that its routing will extract a particular supplier's invoices better; only the team's invoice set can establish that. This is the trade-off I would choose for a weekly shipping cadence because adapter maintenance does not improve the invoice product, while the extraction acceptance rules do.

The boundary matters beyond chat. Image generation and speech should remain optional modules unless the product needs them. In this runtime, automatic speech recognition is currently marked unavailable in the model directory, real-time voice sessions remain pending and region-limited, and there is no dedicated moderation endpoint. A chat model with application-side schema validation can support a moderation workflow, but that is not equivalent to a purpose-built moderation API. Image upscaling is limited to Lanczos. None of those constraints block text invoice extraction, but they should stop a team from pretending one abstraction covers every future media feature equally well.

How should quality beat latency?

Use a gate, not a weighted score that can hide a bad extraction. First require every critical field to meet its acceptance rule: exact invoice number, valid currency, arithmetically consistent totals, and line items that add up within the tolerance your accounting workflow permits. Among models that pass, select the lower-latency result from measurements collected in the deployment region.

I would keep three outcomes for each test invoice: pass, review, and fail. Consider an invoice whose printed total is 37.00, whose two line items also sum to 37.00, and whose purchase-order field is blank. A response with the correct total but a 0.00 purchase order should enter review because the model invented a value for a missing optional field. A response with 37.00 in the subtotal but 73.00 in the total should fail even if every other field is right. The fast response does not get partial credit for a corrupted amount. This decision rule is intentionally boring because customer-support agents need a dependable queue, not an impressive demo, and because explicit validation survives a provider change better than prompt wording does.

Bad totals stop here.

Re-run the set whenever the prompt, parser, model ID, or OCR stage changes. Pin the selected model in production through MODEL_ID; discovery should verify availability, not cause an unreviewed model change on every process start. That separation preserves both flexibility and control.

What I would change at scale

At higher volume, I would add a provider-neutral evaluation record containing a hash of the input, prompt version, model ID, validation outcome, and observed latency. Do not store raw supplier documents in that record by default. Retention and access controls should follow the sensitivity of the invoice data and the obligations that apply to the business.

I would also split synchronous support work from backfills. An agent waiting on one invoice needs a bounded response time and an explicit review path. A historical batch can run asynchronously, checkpoint progress, and tolerate retries. The extraction contract stays identical, which is the useful part of the abstraction.

Finally, keep an exit test. Once per release cycle, run the same fixture set through at least one alternate model. If moving requires a rewrite, the application is already locked in despite the normalized endpoint. If the direct provider consistently wins on a capability that affects revenue or correctness, take the dependency openly. Outsource the undifferentiated plumbing; keep the acceptance rules in your own code.

References

Top comments (0)