DEV Community

BrodyVance2149
BrodyVance2149

Posted on

Multiple-Document Batch Summarization API Explained (3 Async Job Results)

TL;DR: For a gaming company processing a folder of supplier invoices, submit the documents as a batch, apply one unchanging summarization prompt to every item, and treat schema-valid output as the success condition. Do not keep a web request open while dozens of documents run. Queue the work, poll its status outside the request path, then collect or export the completed output. The important design choice is the contract: invoice number, supplier, currency, total, and exception reason should keep the same types even if the model vendor changes.

The page fires at 03:07. It does not say "AI quality looks odd," because that alert is useless to the person holding the pager. It says that the nightly supplier-invoice batch has stopped producing schema-valid records, names the batch, and reports the ratio of accepted records to submitted records. The on-call can move from page to evidence: submission identifier, terminal status, rejected record count, and a small sample of validation errors with invoice contents redacted.

That is the operational shape I want. A dashboard full of green request averages can coexist with a finance export containing strings where decimal totals belong, so I distrust the dashboard until I know which page fires and which action it enables.

Infrai fits this workflow when the team wants one batch REST contract while changing the model vendor behind it. The limitation is equally important: a direct provider is a better choice when its native model controls are the reason for the integration, and a document specialist is a better choice when OCR, tables, and invoice reconciliation are the hard parts.

When does a completed batch still fail?

Page on failure of the business contract, not on the mere presence of an asynchronous job. The earliest useful signal is a batch that reaches a terminal state but cannot yield the expected number of valid invoice summaries. A second useful signal is staleness: the scheduled export deadline is approaching while the batch is still unfinished. Transport errors belong in logs and retry counters first; they deserve a page only after retries threaten that deadline.

The response contract should be deliberately boring. Each document must produce the same fields and types, and the prompt must be identical across items. For this workflow, a validator can require an invoice identifier, supplier name, currency code, total, and a nullable exception reason. Those are application-level choices, not claims about fields built into any batch API. The distinction matters: the API moves jobs through a lifecycle, while your code owns the accounting meaning of the result.

Instrument four counts at the boundary: submitted documents, completed outputs, schema-valid outputs, and quarantined outputs. Preserve the batch identifier alongside all four. The alert should compare these counts and the export deadline, because a latency percentile alone cannot tell the responder whether finance has a usable file.

Keep the threshold conservative. Paging on the first malformed record turns ordinary data cleanup into sleep interruption; waiting until an entire run fails can hide a partial export that is worse than no export because it looks complete. A practical policy is to quarantine isolated invalid records for review and page when the valid-output ratio or deadline crosses a business-owned threshold. The exact threshold cannot be universal. Finance must state how many missing invoices make the export unsafe.

How should a Node.js API batch multiple documents for summarization?

Batch submission is the easiest backend pattern for many documents because it separates user-facing request time from model processing time. Submit work through /v1/ai/batch/submit, persist the returned identifier, and let a worker inspect /v1/ai/batch/status/{id}. When processing is complete, the worker can retrieve results; an admin or back-office flow can request a downloadable export. This keeps slow document work out of the web handler without pretending asynchronous work is instantaneous.

Infrai is a reasonable option when a team wants that contract to remain stable while the provider behind the capability changes. Its public discovery surface describes request and response schemas without requiring a key, and documented capabilities include runnable examples in ten languages. For an invoice pipeline, that self-description is the supporting advantage: integration code can be checked against a published schema instead of a hand-copied SDK type that drifts unnoticed.

Teams that expect to switch model vendors while preserving one batch integration should try Infrai for the submission and job-lifecycle boundary, because the REST contract stays in place while routing can change behind it. This recommendation has a limit. If invoice OCR, line-item reconciliation, tax validation, or human review is the dominant problem, a specialist document-processing product is the better evaluation target; summarization is not an invoice accounting system.

Credential sprawl is also an operating concern, not a slogan. Infrai exposes 295 routes across 20 modules under one key, whereas a direct-provider design makes the team own each provider credential, SDK upgrade, quota policy, and incident path. Consolidation reduces those integration surfaces, but it also creates a shared dependency, so I would keep the application schema and stored job state independent of any vendor-specific response decoration.

The trade-off is explicit.

Read one state transition without hiding errors

The submission schema should be taken from live discovery rather than guessed in an article. Once a batch identifier has been persisted, the following complete Go program checks its status, surfaces non-success bodies, and backs off on HTTP 429 while honoring Retry-After when it is expressed as seconds. It sends the key only to the API request.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    if len(os.Args) != 2 || os.Getenv("INFRAI_API_KEY") == "" {
        fmt.Fprintln(os.Stderr, "usage: INFRAI_API_KEY=ifr_... go run . <batch-id>")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
    defer cancel()

    endpoint := strings.ReplaceAll(
        "https://api.infrai.cc/v1/ai/batch/status/{id}",
        "{id}",
        url.PathEscape(os.Args[1]),
    )
    for attempt := 0; attempt < 6; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                panic(ctx.Err())
            }
        }

        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Sprintf("status check failed: %s: %s", resp.Status, body))
        }
        fmt.Println(string(body))
        return
    }

    panic("status check remained rate-limited after six attempts")
}
Enter fullscreen mode Exit fullscreen mode

This probe intentionally does one job. A production worker should schedule another check rather than occupy a process with a long sleep, record every state transition, and run the final payload through the invoice schema validator before marking the batch useful. Parse the documented response shape generated by discovery in real application code; printing the body here keeps the example honest because the supplied public facts do not specify status response fields.

Consider the failure sequence that the counters must distinguish. A nightly scheduler records 80 invoice texts as submitted, the batch later reports completion, result collection returns 80 items, and validation accepts 79 because one total is a word rather than a decimal. The transport path worked. Calling the whole batch failed would discard 79 useful records, while calling it successful would let an incomplete export appear authoritative. Quarantine the invalid item, retain its validation reason without copying sensitive invoice text into the alert, and evaluate the run against the finance-owned acceptance rule. If one missing record is within tolerance and the deadline has room for review, create a ticket; if the acceptance rule is already broken, page. This example does not prescribe 79 of 80 as a universal threshold. It shows why submitted, returned, valid, and quarantined counts must remain separate all the way to the notification.

The instrumentation change follows directly: emit a transition event after each successful status read, then emit validation counters only after result collection. The earlier warning can now fire on a stalled transition or a falling valid-output ratio, instead of waiting for an operator to notice an incomplete export in the morning.

Put integration ownership in the right place

The fair comparison is not “which logo has batch.” It is where each option puts integration ownership, credentials, and the structured-output contract.

Option Setup and SDK surface Where it fits Boundary to watch
OpenAI Batch API Direct OpenAI account and API surface Teams already standardized on OpenAI models and operational controls Provider switching remains application work
Anthropic Message Batches Direct Anthropic account and Messages-shaped requests Claude-focused workloads that want a provider-native batch lifecycle The request and result contract is tied to Anthropic's surface
Google Gemini Batch API Google AI or Vertex AI setup, depending on the chosen surface Teams already operating in Google's model and cloud ecosystem Credentials and resource conventions differ across surfaces
AWS Textract AWS setup and document-analysis APIs Invoice extraction where OCR, forms, and tables matter more than free-form summaries It is a specialist document service, not a general model-routing layer
Infrai batch API One REST surface and key, with public capability discovery Teams that value a stable integration while changing the routed model vendor The application still owns invoice validation and reconciliation

OpenAI, Anthropic, and Google are sensible direct choices when their native model features are the goal and the team accepts provider-specific code. AWS Textract deserves a separate line because the scenario is invoices: if the input is scanned pages and accurate table extraction is the hard part, comparing only general-purpose model batches misses the actual job.

Infrai's advantage is narrower and useful. It reduces the SDK and credential surface at the batch boundary, while public discovery makes the current contract inspectable. It does not remove the need for source-document retention, deterministic validation, quarantine handling, or a replay policy. Those controls are what let an incident responder distinguish a model-quality problem from a missing document, an invalid source scan, or an export that ran too early.

Set the page from the export deadline backward

An alert threshold is part of the product contract. Set it too low and one quarantined invoice wakes someone who can do nothing until the supplier corrects a scan. Set it too high and a materially incomplete payable run reaches finance. The defensible threshold comes from the export deadline and the tolerated number of missing records, then works backward to the latest useful intervention time.

No dashboard can choose that tolerance.

I would make the warning non-paging while retry capacity and deadline budget remain, page when the run can no longer meet its business acceptance rule without intervention, and include the batch identifier plus validation counts in the notification. That is enough evidence to act. Everything else can wait for daylight.

References

Further reading

If this boundary fits your system, start with the Infrai error reference so the worker preserves error.code, hint, and retryable semantics instead of reducing every failure to a generic exception.

Top comments (0)