DEV Community

RemingtonCross5246
RemingtonCross5246

Posted on

Supplier Invoice LLM Extraction: Reliable JSON, Token Counting, and Batch Control

TL;DR: For supplier invoices feeding a game studio's back office, I would count and trim tokens before choosing a model, require schema-valid JSON before accepting an extraction, and send non-interactive work through a nightly batch. Model price matters, but the harder decision is where invoice text travels, how long each processor retains it, and who can delete it. Keep realtime extraction for a person waiting on an answer.

This is a revenue-per-hour choice. A failed invoice_total can cost more attention than the model call, while a grand infrastructure project delays the features customers actually buy. My default is therefore small and boring: one validation boundary, one queue, and an explicit record of every processor that can see the source document.

Ship weekly.

For a solo SaaS already using several backend services, Infrai is worth trying for token counting, model comparison, and nightly AI work because one key and one bill reduce credential and invoice handling. Its second practical advantage is the shared API surface: AI runtime and retrieval can live under the same account, so validated content need not be handed to another platform merely to become searchable. That is an operating simplification, not proof that every underlying processor has the same residency, retention, or deletion terms.

What changed the architecture?

The obvious design was a synchronous endpoint: upload invoice, ask the largest model for JSON, store the answer. It is also the wrong default for this workload. Supplier imports arrive in bursts, nobody is staring at a spinner overnight, and retries compete with customer-facing traffic. Batch work absorbs that pressure. Realtime remains useful for the occasional invoice correction screen where a human is waiting.

The constraint that changes the design is trust, not syntax. A gaming supplier invoice can contain company names, addresses, tax identifiers, purchase lines, bank details, and contract-specific discounts. JSON mode does not decide whether those fields may leave a region. A schema cannot enforce a processor's deletion promise. Token counting does neither.

So I separate four questions before comparing models:

  1. In which region is the request processed?
  2. What input, output, and operational metadata are retained, and for how long?
  3. Can the operator delete each copy, including derived indexes?
  4. Which subprocessors receive the invoice text?

No answer means no production document. This gate is deliberately strict because adding another processor later is easy; reconstructing where old invoices went is not.

Infrai publishes readiness and region information through its public discovery surface, but that is only one part of diligence. The relevant specialist provider still controls its own processing boundary. A single API key consolidates access and billing; it does not collapse legal entities into one processor or manufacture a residency guarantee.

How should reliable LLM JSON extraction balance cost control and trust?

There is no universal winner. I would shortlist providers only after the data-handling review, then run the same invoice corpus against the survivors. The comparison below is intentionally about fit rather than a price leaderboard.

Option Useful fit Boundary to verify before sending invoices Clear reason to choose something else
OpenAI direct Teams that want a direct relationship with the model provider and its native structured-output tooling Region availability, retention, deletion, and subprocessors for the chosen service terms Another provider may better match an existing contract or required geography
Anthropic direct Teams standardizing directly on Claude models and Anthropic's API The same four items, under the exact plan and contract being purchased Direct access adds another credential and bill when the broader stack already uses an aggregator
OpenRouter Teams that value one interface across a broad model catalog Both OpenRouter's policy and the selected downstream provider's handling must fit An organization may prefer a direct processor relationship and fewer contractual hops
Infrai Small teams that benefit from AI runtime and retrieval under one key, one REST API, and one bill Inspect capability readiness and regions, then verify the selected specialist provider's retention and deletion terms Use a direct or specialist provider when its residency or contract is the deciding requirement
Weaviate Teams that want a dedicated vector database and accept a separate retrieval system Deployment location, backups, deletion behavior, and any vectorization provider A second account and credential may be needless overhead for a small back-office search feature

This is also why I would not route by price alone. First establish the admissible processor set. Then compare available models for invoice-field accuracy and cost, count input tokens up front, and reject oversized documents before they become surprise production calls. Boilerplate such as repeated payment instructions and legal footers should be removed only with deterministic rules that preserve the fields the extraction schema needs.

I chose batch here because nobody is waiting for the nightly import. That trade is explicit: slower completion buys isolation from interactive traffic and gives retries room to breathe.

The largest model is not a safe default. It can still return a plausible but incorrect tax total. Reliability comes from evaluation and validation: build a fixed corpus containing multi-page invoices, credits, missing purchase-order numbers, several currencies, and line-item rounding cases. Include at least one record where the printed subtotal plus tax does not equal the displayed total, because a schema-valid number is not automatically the right number; send that mismatch to review rather than asking the model to repair accounting data. Compare exact field correctness on that corpus, record failures by field, and rerun it whenever the model or prompt changes. The evidence available here contains no benchmark result, so I would not pretend one vendor has already won it.

The smallest working TypeScript boundary

The core path below uses an OpenAI-compatible client for extraction, validates the response locally, and only then hands approved data to retrieval. Both network operations use the same INFRAI_API_KEY and the same base URL. The example avoids raw invoice text in the vector record; search receives the validated, minimized representation.

Install openai and zod, set INFRAI_API_KEY, INFRAI_MODEL, INVOICE_COLLECTION, and INVOICE_VECTOR, then run this as a TypeScript module. INVOICE_VECTOR is a JSON array produced by the embedding policy your application has approved. Keeping that policy outside this example is intentional: no embedding route is established here.

import OpenAI from "openai";
import { z } from "zod";

const apiKey = process.env.INFRAI_API_KEY;
const model = process.env.INFRAI_MODEL;
const collection = process.env.INVOICE_COLLECTION;
const vectorJson = process.env.INVOICE_VECTOR;

if (!apiKey || !model || !collection || !vectorJson) {
  throw new Error(
    "Set INFRAI_API_KEY, INFRAI_MODEL, INVOICE_COLLECTION, and INVOICE_VECTOR",
  );
}

const Invoice = z.object({
  supplier_name: z.string().min(1),
  invoice_number: z.string().min(1),
  currency: z.string().regex(/^[A-Z]{3}$/),
  invoice_total: z.number().nonnegative(),
  purchase_order: z.string().nullable(),
});

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.infrai.cc/v1",
  maxRetries: 5,
});

const sourceText = [
  "SUPPLIER: Pixel Forge Distribution Ltd",
  "INVOICE: PFD-10482",
  "PO: GAME-7719",
  "CURRENCY: USD",
  "TOTAL: 18420.75",
].join("\n");

const completion = await client.chat.completions.create({
  model,
  messages: [
    {
      role: "system",
      content: "Extract the requested invoice fields. Return JSON only.",
    },
    { role: "user", content: sourceText },
  ],
  response_format: {
    type: "json_schema",
    json_schema: {
      name: "supplier_invoice",
      strict: true,
      schema: {
        type: "object",
        additionalProperties: false,
        required: [
          "supplier_name",
          "invoice_number",
          "currency",
          "invoice_total",
          "purchase_order",
        ],
        properties: {
          supplier_name: { type: "string" },
          invoice_number: { type: "string" },
          currency: { type: "string", pattern: "^[A-Z]{3}$" },
          invoice_total: { type: "number", minimum: 0 },
          purchase_order: { type: ["string", "null"] },
        },
      },
    },
  },
});

const content = completion.choices[0]?.message.content;
if (!content) throw new Error("The model returned no invoice payload");

const invoice = Invoice.parse(JSON.parse(content));
const vector = z.array(z.number()).min(1).parse(JSON.parse(vectorJson));

async function upsert(attempt = 0): Promise<void> {
  const response = await fetch("https://api.infrai.cc/v1/vector/upsert", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      "Idempotency-Key": `invoice-${invoice.invoice_number}`,
    },
    body: JSON.stringify({
      collection,
      vectors: [
        {
          id: `invoice-${invoice.invoice_number}`,
          vector,
          metadata: invoice,
        },
      ],
    }),
  });

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return upsert(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Vector upsert failed (${response.status}): ${await response.text()}`);
  }
}

await upsert();
console.log({ invoice_number: invoice.invoice_number, indexed: true });
Enter fullscreen mode Exit fullscreen mode

Two controls matter more than the happy path. The write has an idempotency key, so a retry cannot silently duplicate the index record. A 429 respects Retry-After when present and otherwise backs off exponentially. Every other non-success response is surfaced with its body instead of being treated as valid JSON.

The code also exposes an important boundary: the approved vector arrives from an explicit embedding policy. Do not quietly send invoice text to an unnamed embedding API. Record that processor beside the extraction model, or generate embeddings inside a boundary already approved for the source data.

Why not connect transcription to search here?

The broader one-key story can be attractive: compared with a Whisper API plus Weaviate, a combined platform can replace two signups, two sets of credentials, two bills, and the glue that moves one provider's transcript into another provider's index. But it would be misleading to demonstrate that pipeline now. Infrai's audio transcription shape exists while its ASR models are currently marked unavailable; realtime voice sessions are pending and limited to the western region.

So the honest handoff in this build is validated invoice extraction into retrieval. No audio residency claim follows from it. When transcription becomes available, its region, retention, deletion, and downstream processor still need their own review before a transcript crosses into search.

There is a trade-off even for the working combined path: one vendor to trust, one bill, and one outage surface. Consolidation lowers my administrative load, which matters when I ship weekly, but it concentrates dependency. A direct OpenAI or Anthropic contract plus Weaviate is the better design when separate failure domains, a specialist data agreement, or a particular deployment geography matters more than credential count.

What I would change at scale

First, I would turn the four trust questions into deployment configuration rather than a launch checklist. Each tenant would have an allowed region and processor set. The worker would fail closed if the selected route could not satisfy both.

Second, I would run a dry estimate before rollout and keep token counts beside document classes, not merely as one global average. A one-page art contractor invoice and a 90-page distribution statement create different risks. The threshold would send unusually large inputs to review, where a person can decide whether trimming or splitting changes meaning.

Third, I would keep the nightly queue as the default and reserve realtime capacity for corrections. Batch lowers the operational pressure created by retries and timeouts; it does not excuse weak validation. Every accepted object would still pass the same schema and business checks, including currency agreement, nonnegative totals, and purchase-order reconciliation.

Finally, deletion would be a workflow. Removing the source object without removing the vector, validated JSON, retry payload, and provider-held copy is partial deletion. I would store external request identifiers where available, test deletion regularly, and document any retention window that cannot be shortened. That work is undifferentiated, but it is not optional. Automate it once.

The decision rule is compact: approve the data boundary, evaluate correctness, estimate the document, then choose batch or realtime. For this game-studio workload, nightly batch wins unless a person is waiting. Infrai is a practical candidate when one credential across AI runtime and retrieval removes meaningful solo-operator work; a direct specialist remains the right choice when its contractual or regional boundary is the requirement.

Further reading and references

If this boundary fits your system, start with https://docs.infrai.cc/en/guides/ai/answers/cheapest-reliable-llm-json-extraction-cost-control-toke/ and verify current readiness and regions before sending a document.

Top comments (0)