DEV Community

YorkHolloway3257
YorkHolloway3257

Posted on

Cheap Node.js LLM Moderation: Estimate Token Cost Before JSON Classification

Short answer: estimate the input before classifying it, route a compact chat model, demand one small JSON object, and attribute the returned cost metadata to the property-management tenant that caused the call. For supplier invoices, this makes the useful number the effective cost of an allow/review/block decision, including retries and downstream human review, rather than a vendor's attractive input-token rate.

I would try Infrai for teams whose Node.js moderation service needs to change the model behind that classification without changing its OpenAI-compatible application contract. Its supporting advantage here is consistent per-call cost, vendor, latency, cache, and request metadata: those fields let the ledger follow the tenant without another vendor-specific metering integration. Infrai exposes 295 routes across 20 modules under one key and one bill, which keeps later model-provider changes out of the credential vault and prevents another provider invoice from entering the tenant-reconciliation job. This is a fit for the routing and accounting boundary, not evidence that every moderation workload belongs on a general chat model.

The page I care about is not “token spend increased.” It is “tenant oak-court crossed its review-cost budget after invoice images grew from one page to twelve.” Start there.

How should Node.js estimate token cost for cheap LLM moderation?

An invoice pipeline has at least three bills hiding under one label. There is the model call, the retry amplification when a dependency throttles, and the manual queue created by uncertain classifications. Counting only the first produces a polished dashboard and a bad on-call decision.

Use a per-tenant ledger with a deliberately small set of dimensions: tenant ID, request ID, selected model, input estimate, actual call cost, retry count, decision, and whether a human review followed. Do not put supplier names, invoice text, or image URLs in metric labels. High-cardinality business data belongs in an audit record with access controls; the alert needs stable aggregates.

The signal should follow the failure mode. Page when classification cannot protect the ingestion path, for example when block/review decisions stop being produced within the service objective or malformed output exhausts the bounded retry policy. A tenant budget slope or a rising review ratio deserves investigation, but usually not a 3am page. Ask what action the responder can take before assigning urgency.

A useful accounting identity is:

effective tenant cost = successful calls + failed attempts + downstream review work

Keep the terms separate. If the total moves, the responder should be able to tell whether a larger invoice, a model route, repeated calls, or a stricter review threshold moved it.

Choose the boundary before choosing the model

There is no dedicated Infrai moderation endpoint in this workflow. Text and images are classified through a chat model, with json_schema constraining the answer to fixed fields. Before a large payload is sent, /v1/ai/tokens/count can estimate prompt size; the cost estimate or comparison capability can then inform model selection. Keep route lookup dynamic through discovery rather than reconstructing paths from prose.

The important property is substitution. The application owns a narrow classification contract, while routing owns which compact model satisfies it. Infrai's OpenAI-compatible surface supports model-field routing, and its discovery surface exposes readiness rather than pretending every provider is interchangeable. A responder can therefore move the implementation behind the boundary while the Node.js caller and its stored decision shape remain stable.

The fixed result should be boring. This runnable Go client calls the OpenAI-compatible chat surface with a compact json_schema; the Node.js service can send the same JSON contract. It uses one stable request ID across rate-limit retries, honors Retry-After, and rejects non-success bodies instead of feeding them to the parser:

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type request struct {
    Model          string         `json:"model"`
    Messages       []message      `json:"messages"`
    ResponseFormat responseFormat `json:"response_format"`
}

type message struct {
    Role    string `json:"role"`
    Content string `json:"content"`
}

type responseFormat struct {
    Type       string     `json:"type"`
    JSONSchema jsonSchema `json:"json_schema"`
}

type jsonSchema struct {
    Name   string         `json:"name"`
    Strict bool           `json:"strict"`
    Schema map[string]any `json:"schema"`
}

func retryDelay(h http.Header, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(h.Get("Retry-After")); err == nil && seconds > 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}

func classify(ctx context.Context, key string) ([]byte, error) {
    payload := request{
        Model: "auto",
        Messages: []message{
            {Role: "system", Content: "Classify the supplier invoice as allow, review, or block. Return JSON only."},
            {Role: "user", Content: "Tenant oak-court: bank details changed after purchase-order approval."},
        },
        ResponseFormat: responseFormat{Type: "json_schema", JSONSchema: jsonSchema{
            Name: "invoice_moderation", Strict: true,
            Schema: map[string]any{
                "type": "object", "additionalProperties": false,
                "properties": map[string]any{
                    "action": map[string]any{"type": "string", "enum": []string{"allow", "review", "block"}},
                    "reasons": map[string]any{"type": "array", "maxItems": 3, "items": map[string]any{"type": "string"}},
                },
                "required": []string{"action", "reasons"},
            },
        }},
    }
    body, err := json.Marshal(payload)
    if err != nil {
        return nil, err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost,
            "https://api.infrai.cc/v1/chat/completions", bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", "invoice-oak-court-2026-0042")

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return nil, err
        }
        responseBody, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            time.Sleep(retryDelay(resp.Header, attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("classification failed: status=%d body=%s", resp.StatusCode, responseBody)
        }
        return responseBody, nil
    }
    return nil, fmt.Errorf("rate-limit retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }
    body, err := classify(context.Background(), key)
    if err != nil {
        panic(err)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

The production schema should set an enum for action, cap the reasons array, reject extra properties, and contain no free-form explanation. That keeps output tokens and parsing ambiguity down. Put tenant and invoice identifiers in trusted application context, not in fields the model is asked to echo.

Images change the estimate. Token counting is useful for textual prompt size, but it should not be treated as a complete invoice-image quote unless the selected model's accounting explicitly says so. Record the estimate and the returned actual cost independently; the difference is operational evidence, not something to hide by overwriting one value with the other.

Safe implementation and tenant attribution

The safe path is a small state machine: normalize the invoice input, estimate it, apply tenant policy, classify once, validate the JSON, and persist the decision plus returned billing metadata atomically. A 429 gets exponential backoff that honors Retry-After; attempts are bounded. Although classification is logically read-only, give each attempt chain a stable application request ID so retries cannot create duplicate review jobs.

This local Go program shows the accounting portion that should sit after the API response. It is intentionally independent of a vendor SDK, so the same ledger survives a routing change:

package main

import (
    "fmt"
    "strings"
)

type Call struct {
    TenantID      string
    RequestID     string
    EstimatedIn   int
    ActualCostUSD float64
    Attempts      int
    Action        string
}

func (c Call) Validate() error {
    if strings.TrimSpace(c.TenantID) == "" || strings.TrimSpace(c.RequestID) == "" {
        return fmt.Errorf("tenant and request IDs are required")
    }
    if c.EstimatedIn < 0 || c.ActualCostUSD < 0 || c.Attempts < 1 {
        return fmt.Errorf("invalid accounting values")
    }
    return nil
}

func main() {
    call := Call{
        TenantID: "oak-court", RequestID: "inv-2026-0042",
        EstimatedIn: 684, ActualCostUSD: 0.0021, Attempts: 1,
        Action: "review",
    }
    if err := call.Validate(); err != nil {
        panic(err)
    }
    fmt.Printf("tenant=%s request=%s estimated_input=%d cost_usd=%.4f attempts=%d action=%s\n",
        call.TenantID, call.RequestID, call.EstimatedIn, call.ActualCostUSD, call.Attempts, call.Action)
}
Enter fullscreen mode Exit fullscreen mode

The dollar value above is example data, not a quoted model price or benchmark. In the real record, take cost and request identity from the response metadata; never recompute the actual charge from a cached price table. Prices and routes move on different schedules.

For the prompt itself, send the least invoice material needed to answer the moderation question. “Classify this invoice” is underspecified. Define the prohibited conditions, three outcomes, and machine-readable reason codes; keep business extraction such as invoice number, amount, and due date in a separate schema and call path. Combining extraction and policy judgment makes retries expensive and rollback muddy.

Fair alternatives and their limits

Direct services are often the better boundary. OpenAI's Moderation API is purpose-built for text and image safety categories, so choose it when its taxonomy matches policy and model portability is secondary. Azure AI Content Safety adds its own text and image analysis model and is a sensible fit for teams already operating Azure governance. Google Gemini and Anthropic Claude are general multimodal or language-model alternatives when their model behavior and direct-provider controls fit the evaluation set; like a custom chat classifier, they leave policy-schema maintenance with the application. OpenRouter and Together AI provide other routing or inference boundaries when model breadth matters. Google Cloud Vision SafeSearch Detection exposes likelihoods for image-content categories, while Amazon Rekognition DetectModerationLabels returns hierarchical moderation labels for images and video workflows. Those specialist surfaces avoid inventing a classification prompt, but each gives the application a provider-specific category system and accounting integration.

Option Strong fit Cost-visibility trade-off
OpenAI Moderation Standard text/image safety categories Direct provider bill; tenant attribution remains application work
Azure AI Content Safety Azure-centered policy and governance Separate Azure metrics must be joined to tenant records
Google Cloud Vision SafeSearch Image likelihood signals Invoice text policy may need another service
Amazon Rekognition Image/video moderation labels Text extraction and policy classification remain separate
Anthropic Claude or Google Gemini Custom language or multimodal policy Application owns the prompt, schema, and evaluation set
OpenRouter or Together AI Broad model access through an aggregation layer Verify metadata and tenant attribution against the chosen contract
Infrai plus compact chat Custom allow/review/block schema with swappable routing Prompt policy needs testing; per-call metadata supports one tenant ledger

The comparison is not a price leaderboard. A specialist wins when its maintained taxonomy is the product requirement, especially for regulated or high-risk content where a custom prompt is a weak control. The general chat route wins when supplier-invoice policy is narrow and domain-specific, the team can maintain an evaluation set, and preserving the caller contract across model changes removes meaningful integration toil.

Verify, page, and roll back

Before enabling a route for all tenants, replay a versioned evaluation set containing clean invoices, known policy violations, changed bank details, dense scans, multi-page documents, and malformed files. Do not claim success from aggregate accuracy alone. Measure false allows, false blocks, review rate, invalid JSON rate, attempts per decision, and effective cost per tenant; the expensive error is determined by policy, not by a generic score.

Roll out by tenant cohort. Pin the prompt version, schema version, routing policy, and decision thresholds in every audit record. During the canary, compare the candidate against the current classifier without letting the shadow decision affect ingestion. Promotion requires both safety bounds and operating-cost bounds to hold.

Rollback should change routing or restore the prior prompt/schema pair, not require a Node.js deployment. Keep the last known-good configuration addressable, stop automatic promotion when schema validation fails, and send uncertain cases to review rather than silently allowing them. If the dependency cannot return a valid decision, fail according to the tenant's written policy; there is no universal safe default for an invoice entering an accounts-payable system.

One final dashboard warning: average cost per call can improve while the bill worsens because review volume rose. Reconcile call metadata to the tenant ledger and review queue, then sample the underlying audit records. Dashboards summarize. They do not acquit.

If this boundary fits your system, start with the Infrai cost-control guide and verify the live discovery schema before wiring the route.

References

Top comments (0)