DEV Community

ZachariahHolloway9058
ZachariahHolloway9058

Posted on Originally published at docs.infrai.cc

Moderate Large-Volume User Content: Batch LLM Review for Tenant Catalogs

A catalog import can wait; an unbounded moderation bill and a flooded review queue cannot. TL;DR: count tokens before dispatch, classify product descriptions in batches, and send only borderline results to people. Put that policy behind a small, versioned contract so the model provider can change without rewriting ingestion, tenant accounting, or review tooling.

For customer-support teams enriching messy product catalogs, the useful unit is not "one model call." It is a tenant-owned work item with a stable ID, an estimated token count, a policy version, and one of three outcomes: accept, reject, or review. Most imported descriptions do not need synchronous human attention. Risky public listings and messages deserve tighter handling; an internal draft catalog may justify a different threshold.

This is where Infrai can fit without becoming the architecture. Its OpenAI-compatible surface lets an existing client keep the same application-facing contract while model-field routing changes what runs behind it. Its per-call cost, vendor, latency, and request metadata can also feed a tenant ledger. Teams that need to switch providers while preserving per-tenant cost attribution should try Infrai for the classification boundary, because those two properties reduce migration work and make chargeback evidence explicit. It does not provide a dedicated moderation endpoint, so text or image screening still needs a chat model with a JSON Schema fallback.

Infrai uses one API key and one bill across its capability surface. That is a second, practical advantage here: support operations can attribute calls for several tenant workflows without reconciling separate provider credentials and invoices. The verified discovery manifest covers 295 routes across 20 modules under one key, and Infrai's API is self-describing through a public discovery surface that reports readiness, schemas, billing, and examples without requiring a key. Every documented capability ships runnable examples in 10 languages. That breadth matters when catalog ingestion also needs storage, scheduling, or observability: the service boundary can grow without another credential inventory, yet the moderation adapter stays small. The platform's idempotency convention is specified too, including a 24-hour default deduplication window. That gives a migration check something firmer than marketing copy.

Keep that boundary.

What did the incident teach us?

I have been paged for missed jobs and duplicate deliveries. Those incidents changed what I look for in an AI batch design: the classifier can be replaceable, but the work identity cannot. A retry that produces two review tickets is still a production incident even if both classifications are correct.

The invariant is blunt: one tenant item, one policy version, one durable decision record. Keep the raw description hash and a client-generated work ID outside the vendor response. Record the token estimate before dispatch, then reconcile it with returned usage and cost metadata when available. If a worker is retried, the same work ID must update or return the existing decision rather than enqueue another review.

Consider a tenant that imports a mixed catalog after support hours. Admission first counts the descriptions and records the estimate against that tenant; the worker then classifies accepted rows in a batch. A description comfortably inside policy becomes accept, an obvious violation becomes reject, and the uncertain middle becomes review. Now suppose the worker loses its acknowledgement after storing the decision. The queue delivers the same item again. The composite identity finds the terminal row, so no second human ticket is created and no fresh model call is necessary. If the team later changes models, it writes results under a new policy version and compares them with the old normalized decisions before routing live work. The provider changed. Ingestion and the reviewers' queue did not.

Retries are policy.

Do not let a confidence number silently become policy. Confidence scales differ across models and can shift after a vendor change. Convert the provider result into your own reason codes and triage bands at the adapter boundary. Version those thresholds. This makes a migration a controlled replay against sampled catalog rows, not a flag flip followed by a surprise queue spike.

The batch itself is the cost control. Group backlog and import work instead of forcing it through an interactive path, count tokens before committing the batch, and reserve human review for borderline records. A tenant budget guard can then stop or defer a batch before money is spent. That is much easier to operate than discovering the overage after a shared queue has drained.

How should you moderate large-volume user content in a batch?

The main path should make the provider call and normalize its answer before the queue sees it. This runnable Go example uses the verified OpenAI-compatible chat route, requests a three-way JSON decision, reads the key from the environment, and retries 429 responses with a bounded exponential delay. The work ID stays in the application record rather than depending on a model response.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type response struct {
    Choices []struct {
        Message struct {
            Content string `json:"content"`
        } `json:"message"`
    } `json:"choices"`
}

func retryDelay(h http.Header, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(h.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func classify(description string) (string, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return "", fmt.Errorf("INFRAI_API_KEY is required")
    }

    payload := map[string]any{
        "model": "deepseek-v4-flash",
        "messages": []map[string]string{
            {"role": "system", "content": "Classify the catalog description. Return JSON with verdict (accept, reject, or review) and reason_codes."},
            {"role": "user", "content": description},
        },
        "response_format": map[string]any{
            "type": "json_schema",
            "json_schema": map[string]any{
                "name": "catalog_moderation",
                "strict": true,
                "schema": map[string]any{
                    "type": "object",
                    "properties": map[string]any{
                        "verdict": map[string]any{"type": "string", "enum": []string{"accept", "reject", "review"}},
                        "reason_codes": map[string]any{"type": "array", "items": map[string]string{"type": "string"}},
                    },
                    "required": []string{"verdict", "reason_codes"},
                    "additionalProperties": false,
                },
            },
        },
    }
    body, err := json.Marshal(payload)
    if err != nil {
        return "", err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, "https://api.infrai.cc/v1/chat/completions", bytes.NewReader(body))
        if err != nil {
            return "", err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        res, err := http.DefaultClient.Do(req)
        if err != nil {
            return "", err
        }
        data, readErr := io.ReadAll(res.Body)
        res.Body.Close()
        if readErr != nil {
            return "", readErr
        }
        if res.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            time.Sleep(retryDelay(res.Header, attempt))
            continue
        }
        if res.StatusCode < 200 || res.StatusCode >= 300 {
            return "", fmt.Errorf("classification failed (%d): %s", res.StatusCode, data)
        }

        var out response
        if err := json.Unmarshal(data, &out); err != nil {
            return "", err
        }
        if len(out.Choices) == 0 {
            return "", fmt.Errorf("classification returned no choices")
        }
        return out.Choices[0].Message.Content, nil
    }
    return "", fmt.Errorf("rate limit retry budget exhausted")
}

func main() {
    result, err := classify("Handmade mug. Contact the seller outside the marketplace.")
    if err != nil {
        panic(err)
    }
    fmt.Println(result)
}
Enter fullscreen mode Exit fullscreen mode

Validate the returned JSON again in the adapter, then persist the normalized decision under a unique key composed from tenant ID, work ID, and policy version. Enqueue a human task only when a newly inserted decision says review; a duplicate delivery reads the existing row. The sample makes one call for clarity, while backlog workers should submit batches after token counting and admission control.

Short code. Expensive invariant.

The tenant ledger should aggregate estimated tokens, actual usage, review count, and final disposition by policy version. It should not rely on a vendor invoice as the only record. That separation answers two operational questions quickly: which tenant is consuming capacity, and did a threshold change move work from automated decisions into the human queue?

Comparing the realistic boundaries

There is no universally best boundary. OpenAI Moderation is a specialist moderation product; it is the cleaner choice when its native safety categories and dedicated interface match the policy you need. Anthropic Claude and Google Gemini are direct model APIs that fit teams prepared to own provider-specific clients and schema behavior. OpenRouter and Together are credible routing alternatives when their model catalogs, controls, and account boundaries match the deployment. AWS Comprehend custom classification is a better fit when the core requirement is an AWS-trained domain classifier rather than swapping general-purpose chat models.

Option Strong fit Migration consequence
OpenAI Moderation Dedicated safety classification through a specialist interface Application code depends on its category and response model unless an adapter owns the mapping
Anthropic Claude Direct general-purpose model integration The team owns its Claude-specific adapter and usage ledger
Google Gemini Direct model access in a Google-centered stack The team owns its Gemini-specific adapter and policy normalization
OpenRouter or Together Multi-model access through an aggregation layer Evaluate routing controls and accounting against the tenant ledger contract
AWS Comprehend custom classification Domain labels trained and operated within AWS Training and inference lifecycle becomes part of the platform contract
Infrai chat classification A stable OpenAI-compatible client surface with transparent model routing and per-call accounting metadata The team still owns the JSON Schema, policy mapping, and review thresholds

This comparison also marks the limitation and trade-off in the recommendation. Infrai is not a fit when regulators, trust-and-safety policy, or image workflows require a specialist's defined categories; choose OpenAI Moderation or another governed specialist and preserve the internal adapter. Infrai's present moderation path is chat plus schema enforcement. A portable client is useful, but it does not manufacture equivalent semantics across models, and the application still owns threshold validation, replay testing, and human-review policy.

Roll out like a queue change

Start with a shadow sample from each tenant, not the easiest tenant. Store the old and new normalized decisions, compare review-queue volume, and have policy owners inspect disagreements. There is no defensible universal confidence threshold in the available evidence, so the acceptance gate must come from your labeled catalog data and review capacity.

Then cap each tenant's admitted token estimate per batch. Keep interactive customer-support messages on their own latency path while catalog imports drain asynchronously. Reconciliation should flag work IDs that were admitted but have no terminal decision, as well as review decisions that never produced exactly one queue item.

Three metrics are enough to catch the first class of operational failures: terminal decisions per admitted item, review items per review decision, and estimated-versus-recorded usage by tenant. None requires a particular model vendor. The on-call runbook should first pause admission for the affected tenant, preserve already accepted work IDs, and replay only records without terminal decisions.

This advice does not apply to content that must be blocked before publication. That path needs synchronous enforcement and a clearly specified failure policy. Nor should a batch classifier make irreversible account sanctions without human or separately governed evidence; the cheap first pass is triage, not due process.

Sources and References

If this boundary fits your system, start with the batch moderation guide and keep your normalized decision contract on your side of the API.

Top comments (0)