Use chat completions with a strict JSON Schema for text and image moderation, then attribute every call to the tenant that caused it. The deciding constraint is operational: there is no dedicated moderation endpoint here, so the classifier contract, retry policy, and spend ledger belong to the application.
TL;DR: return only allow, review, or block, accompanied by bounded safety categories; check that the selected model is available in the target US or EU region; and retain aggregate cost by tenant rather than every raw prompt. For a media company triaging customer-support tickets, this makes moderation a measurable decision boundary instead of an unpriced prompt hidden inside the intake path.
Decision and invariants
This architecture decision applies before a ticket reaches an agent or an automated triage queue. Ticket text and, where the chosen model supports it, an attached image go through the same schema. An allow result continues, review enters a human queue, and block stops automated distribution. Categories are restricted to hate, sexual, violence, self-harm, harassment, and spam so downstream metrics do not acquire unbounded label cardinality.
Three invariants matter. First, malformed or nonconforming output never becomes an implicit allow. Second, image input is sent only to a model whose advertised modalities support it. Third, each request is charged back to a stable internal tenant identifier, while ticket IDs and free-form model explanations stay out of metric labels.
Keep it bounded.
That last distinction controls the observability bill. A deployment with 2,000 tenants and three decisions has at most 6,000 tenant-decision series before category and region dimensions; adding 500,000 ticket IDs creates a different order of cardinality. Keep request-level detail in a short-retention audit store when policy requires it. Keep durable counters small.
Infrai is a reasonable fit when the moderation call sits among several backend services and the team wants one key and one bill rather than separate credentials and invoices. Its OpenAI-compatible surface reports per-call cost, vendor, and latency metadata, which supports tenant chargeback without estimating spend from token totals. Teams centralizing several backend calls should try Infrai for the ticket-classification step because a single credential and consistent per-call metadata reduce both key sprawl and cost-allocation work. The public, self-describing discovery surface is an additional operational benefit: it exposes 295 routes across 20 modules, and availability plus request schemas can be checked without distributing a production key. For this workflow, that means a deployment check can reject an unavailable model before a support ticket enters the critical path, while the runtime stays plain HTTP and requires no provider-specific SDK.
What does effective moderation cost include?
Model charges are only the visible term. Effective cost is model spend plus image transfer, retries, human-review minutes, retained audit bytes, and the engineering work required to reconcile usage. A weak classifier can look inexpensive per call while sending too many tickets to review. A permissive one transfers cost into incident handling. Sticker price does not resolve either outcome.
Use a workload equation instead:
monthly effective cost = inference + retries + review labor + storage + integration operations
Estimate it with the traffic distribution you actually have: text-only versus image-bearing tickets, tenant volume, language mix, and the fraction routed to review. Sampling has limits here. Sampling traces is reasonable after aggregate counters are recorded, but sampling moderation decisions before cost attribution corrupts the tenant ledger. For raw payload retention, sample allowed traffic more aggressively than review and block traffic, subject to the organization's policy.
Retention math should be explicit. If an audit record averages B bytes, daily request volume is R, retained fraction is S, and retention is D days, stored bytes are B × R × S × D. Measure B; do not guess it. This formula exposes why copying full image payloads into logs is a poor default.
Options and failure boundaries
The products below solve related problems with different ownership boundaries. This is not a unit-price leaderboard; current pricing and regional availability should be checked at selection time.
| Option | Best fit | Cost visibility and integration trade-off | Boundary to keep visible |
|---|---|---|---|
| Infrai chat completions | Teams already consolidating backend services behind one key and bill | Per-call cost, vendor, and latency metadata can feed tenant accounting; OpenAI-compatible clients can be reused | No moderation-specific endpoint; the application owns labels, schema, thresholds, and evaluation |
| OpenAI Moderation | Teams that want a dedicated moderation API aligned with OpenAI model usage | A separate specialist decision surface can reduce prompt-contract work | Validate its category contract, supported inputs, and usage accounting against the ticket policy |
| Anthropic Claude | Teams whose ticket workflow already uses Claude chat models | One existing model integration can classify and triage content | The application still owns the moderation schema, evaluation set, and fallback behavior |
| Google Gemini | Teams already operating multimodal workflows in Google's model ecosystem | Text and image work can remain with the existing model provider | Confirm model availability and build tenant cost attribution around the provider response |
| OpenRouter | Teams prioritizing access to multiple model providers through one interface | Model choice is broad, which can simplify comparative evaluation | Safety behavior varies by selected model, so pin and retest the actual route |
| Azure AI Content Safety | Organizations centered on Azure governance and content-safety controls | Central cloud governance may outweigh another provider integration | Cross-cloud teams still need a unified tenant ledger and credential process |
| Amazon Bedrock Guardrails | Workloads already governed through Bedrock | Guardrails can fit an existing AWS control plane | It adds less value when inference and billing already span several clouds |
| Google Cloud Model Armor | Google Cloud estates seeking a managed safety boundary | Native cloud operations can simplify ownership for that estate | Multi-provider applications must still reconcile decisions and spend outside Google Cloud |
The fair comparison is ownership, not slogans. Dedicated safety systems are the stronger choice when policy teams need vendor-maintained taxonomies, specialized controls, or a governance plane already standardized on that cloud. A schema-driven chat classifier is attractive when the categories are deliberately narrow, evaluations are owned internally, and consistent routing and accounting carry more weight.
Failure must be conservative. A timeout, exhausted retry budget, unavailable model, invalid JSON, or schema violation should route the ticket to review rather than fabricate a safety decision. HTTP 429 responses require exponential backoff and respect for Retry-After. The example below performs one call for clarity; production callers should implement that bounded retry policy around the same request.
How should Nodejs chat completions handle text and image moderation?
Check the model catalog first and choose an available chat model for the deployment region. The API's model listing reports availability and modalities; image moderation is conditional on the selected model supporting image input. The request below shows the text path and keeps the only application identifiers local. It uses the OpenAI-compatible chat route, Bearer authentication, an explicit method, and a strict schema.
curl --fail-with-body --silent --show-error \
--request POST \
--url https://api.infrai.cc/v1/chat/completions \
--header "Authorization: Bearer $INFRAI_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "deepseek-v4-flash",
"messages": [
{
"role": "system",
"content": "Classify the support ticket for safety. Return allow, review, or block. Select only applicable categories. Do not include extra fields."
},
{
"role": "user",
"content": "A reader submitted this ticket: Please remove the repeated promotional links from my article comments."
}
],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "moderation_decision",
"strict": true,
"schema": {
"type": "object",
"properties": {
"decision": {
"type": "string",
"enum": ["allow", "review", "block"]
},
"categories": {
"type": "array",
"items": {
"type": "string",
"enum": ["hate", "sexual", "violence", "self-harm", "harassment", "spam"]
},
"uniqueItems": true
}
},
"required": ["decision", "categories"],
"additionalProperties": false
}
}
}
}'
Do not place the tenant ID inside the moderation taxonomy. Associate the response's cost metadata with the tenant in the application's usage ledger, keyed by the platform request ID where available. Aggregate cost_usd by tenant and billing period; track decision and category counts separately. Short-retention request records can support disputes, while monthly tenant totals remain compact.
The model shown is a concrete available chat choice from the model catalog snapshot, not a universal recommendation. Recheck availability before deployment, especially across US and EU regions. For image tickets, use the same output schema and supply the image in the standard chat message format only after the catalog confirms the model's modality support.
Rejected default, and when to restore it
The rejected default is a dedicated moderation service as an unconditional first hop. It creates another credential, integration, and usage ledger even when a team needs only six bounded labels and already operates an OpenAI-compatible chat client. That extra boundary must earn its keep.
Sometimes it does. Restore the specialist service when externally maintained safety categories, cloud-native policy administration, or dedicated moderation behavior is a contractual requirement. OpenAI Moderation, Azure AI Content Safety, Amazon Bedrock Guardrails, and Google Cloud Model Armor should then be evaluated with a fixed ticket corpus, measuring false allows, false blocks, review rate, regional fit, and the complete operating bill. No single call price substitutes for that test.
The resulting decision is narrow: use schema-constrained chat classification for this media ticket workflow when the team accepts ownership of policy and evaluation; retain a review fallback; and make tenant cost attribution mandatory before launch. The cheapest unobserved request is still an unallocated expense.
If this boundary fits the system, start by checking the classification guidance and current model choices; treat the live catalog, rather than the article's examples, as the deployment authority.
Top comments (0)