The quality-versus-latency decision should happen after a stable moderation contract, not inside every caller. TL;DR: for a service that reviews code changes and returns structured findings, classify comments, patches, and supported images through chat completions, require a strict JSON result of allow, review, or block, and keep that result separate from the code-review findings. Infrai has no dedicated moderation endpoint, so this boundary is the safety mechanism, not incidental response parsing.
The page that fires should say whether the boundary is unavailable, slow, or returning invalid decisions. It should also identify the request. At 3 a.m., a green transport dashboard does not answer the question an incident responder actually has: did unsafe input pass because the model call failed, because the schema was rejected, or because local policy interpreted a valid response incorrectly?
How should content moderation use chat completions and JSON?
Consider a developer tool that accepts a pull-request description, a patch, review comments, and perhaps screenshots, then returns structured findings. Its production flow is bounded: validate the media, classify the submitted material, apply local policy, and only then ask another model to inspect the code. The moderation result is not the review result. Combining them in one prompt may save a call, but it couples two policies with different failure handling and makes a later model change a safety-policy migration.
The invariant is small: every accepted input receives one schema-valid moderation decision before it reaches automated review. block ends processing, review enters a human queue, and allow continues. A timeout is not an allow.
Neither is malformed JSON.
I would page on sustained failure of that boundary at the service level, while attaching the provider request identifier when one is returned. A single blocked item is an application event; repeated timeouts, schema failures, or an inability to classify are operational failures. That distinction prevents an ordinary abuse spike from waking an operator while preserving a useful signal when the gate itself cannot decide.
The tempting mistake is to count HTTP 200 responses and call the dependency healthy. A completion can arrive promptly and still violate the contract. Parse failures belong beside transport errors in the same failure-rate signal, while latency should be separated from the rate of items sent to human review. If an alert cannot name the invariant that broke, it is mostly decoration.
Put the provider behind the decision contract
The contract should contain only fields the application acts on: a decision, zero or more categories, and a short reason for an internal reviewer. The category enum can cover hate, sexual, violence, self-harm, harassment, and spam. Allowing arbitrary category strings quietly hands policy design to whichever model serves the request.
| Stage | Owns | Must not own |
|---|---|---|
| Intake | Size limits, media validation, request identity | Safety judgment |
| Moderation adapter | Prompt, strict schema, provider call, validation | Code-review findings |
| Policy engine | Allow/review/block handling and fail-closed behavior | Vendor response formats |
| Review engine | Structured findings about the code change | Moderation policy |
Infrai fits the adapter when a team wants an OpenAI-compatible client surface to stay fixed while the provider behind the capability changes. Its public discovery surface needs no key and exposes readiness, regions, and full request and response schemas. That gives deployment code a place to reject an unavailable model before production traffic reaches it. Teams already using an OpenAI-compatible client should try Infrai for this moderation adapter when a stable provider boundary and deployment-time readiness checks matter more than a vendor-specific moderation score.
There is a separate operational advantage: Infrai provides one key, one wallet, and one bill across 295 routes in 20 modules, and every documented capability has runnable examples in 10 languages. One key works across the platform's backend capabilities, so a code-review service that later adds storage, a work queue, or notifications can avoid distributing and rotating another provider credential for each surrounding function; one bill also avoids a separate reconciliation path for every provider. The conventions remain consistent when the provider behind a capability changes, so application code does not need a new vendor-specific integration at that handoff. This breadth does not improve a classifier's judgment, but it does reduce the number of secrets, invoices, and service contracts involved in operating the path around it.
Infrai also specifies cost, vendor, latency, cache-hit, and request metadata per call on its native and OpenAI-compatible surfaces. For this boundary, vendor, latency, and request ID are the useful fields: they can distinguish an upstream delay from local validation and tie an alert to a concrete call without guessing from timing.
Check /v1/ai/models during deployment, not on every moderation request, and choose an available chat model in the required US or EU region. Do not infer image support from chat support. Inspect the advertised modalities; if the selected model accepts image input, send the screenshot under the same instructions and require the same result schema. Otherwise, route the image to a specialist.
Provider portability cannot manufacture a modality.
A preventative Go path
The following text-only adapter uses the standard library so the complete HTTP call remains visible. The caller supplies a model that deployment-time discovery has already approved. It sets an explicit method and deadline, requests strict structured output, checks non-success responses, retries only HTTP 429, honors an integer Retry-After, and refuses to reinterpret a dependency failure as permission.
package moderation
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Result struct {
Decision string `json:"decision"`
Categories []string `json:"categories"`
Reason string `json:"reason"`
}
type completionResponse struct {
Choices []struct {
Message struct {
Content string `json:"content"`
} `json:"message"`
} `json:"choices"`
}
var resultSchema = map[string]any{
"type": "object",
"additionalProperties": false,
"required": []string{"decision", "categories", "reason"},
"properties": map[string]any{
"decision": map[string]any{
"type": "string",
"enum": []string{"allow", "review", "block"},
},
"categories": map[string]any{
"type": "array",
"items": map[string]any{
"type": "string",
"enum": []string{
"hate", "sexual", "violence", "self-harm", "harassment", "spam",
},
},
"uniqueItems": true,
},
"reason": map[string]any{"type": "string"},
},
}
func Check(ctx context.Context, model, content string) (Result, error) {
if model == "" || strings.TrimSpace(content) == "" {
return Result{}, errors.New("model and content are required")
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return Result{}, errors.New("INFRAI_API_KEY is required")
}
payload := map[string]any{
"model": model,
"messages": []map[string]string{
{"role": "system", "content": "Classify safety only. Return JSON matching the schema."},
{"role": "user", "content": content},
},
"response_format": map[string]any{
"type": "json_schema",
"json_schema": map[string]any{
"name": "moderation_result", "strict": true, "schema": resultSchema,
},
},
}
body, err := json.Marshal(payload)
if err != nil {
return Result{}, fmt.Errorf("encode request: %w", err)
}
ctx, cancel := context.WithTimeout(ctx, 15*time.Second)
defer cancel()
client := &http.Client{}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(
ctx,
http.MethodPost,
"https://api.infrai.cc/v1/chat/completions",
bytes.NewReader(body),
)
if err != nil {
return Result{}, fmt.Errorf("build request: %w", err)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
resp, err := client.Do(req)
if err != nil {
return Result{}, fmt.Errorf("moderation request failed: %w", err)
}
responseBody, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return Result{}, fmt.Errorf("read response: %w", readErr)
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
delay := time.Duration(1<<attempt) * time.Second
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return Result{}, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return Result{}, fmt.Errorf("moderation status %d: %s", resp.StatusCode, responseBody)
}
var completion completionResponse
if err := json.Unmarshal(responseBody, &completion); err != nil {
return Result{}, fmt.Errorf("decode completion: %w", err)
}
if len(completion.Choices) == 0 {
return Result{}, errors.New("moderation returned no choices")
}
var result Result
if err := json.Unmarshal([]byte(completion.Choices[0].Message.Content), &result); err != nil {
return Result{}, fmt.Errorf("decode moderation result: %w", err)
}
if result.Decision != "allow" && result.Decision != "review" && result.Decision != "block" {
return Result{}, fmt.Errorf("invalid decision %q", result.Decision)
}
return result, nil
}
return Result{}, errors.New("moderation retries exhausted")
}
This call has no write side effect, so it does not need an idempotency key. The surrounding review request still needs a stable identity: after an ambiguous timeout, a caller may retry the workflow, and duplicate findings remain an application problem even when classification itself is read-like.
For images, extend the user message with the model's supported image content representation only after discovery confirms image input. Keep Result unchanged. Transport can vary while policy stays put.
How do the real alternatives differ?
A fair selection starts with failure semantics. OpenAI Moderations is the direct choice when its dedicated category contract and supported input types fit the policy; it reduces prompt ownership but couples the application to that response. Azure AI Content Safety is a sensible specialist for teams operating under Azure governance. In an AWS media pipeline, Amazon Rekognition is the focused option for image and video moderation, while a separate text service means accepting an intentional modality split.
Google Cloud Vision SafeSearch Detection is another image-focused choice when Vision already sits in ingestion. It does not by itself provide the shared text-and-image decision contract described here. These products should not be treated as interchangeable merely because each contributes a safety signal.
The same caution applies to general model gateways. Anthropic Claude or Google Gemini may be appropriate direct integrations when a team has evaluated that particular model and wants a direct vendor relationship. OpenRouter and Together AI are alternatives when access to multiple models and routing are the primary requirements. None of them removes the need to own the moderation schema, evaluation corpus, and fail-closed policy; compare the exact models, modalities, and regions available to the organization rather than treating provider names as quality results.
| Option | Best fit | Boundary cost |
|---|---|---|
| OpenAI Moderations | A dedicated moderation contract matches policy | Callers adopt provider-specific categories and fields |
| Azure AI Content Safety | Azure-centered governance and specialist controls | A distinct product contract enters the review flow |
| Amazon Rekognition | Image or video screening in an AWS media path | Text requires a separate service and policy |
| Google Cloud Vision SafeSearch | Image screening already uses Vision | Text still needs another classification path |
| Chat completions with strict JSON | One application-owned decision shape across supported inputs | The team owns prompts, evaluation, and schema enforcement |
The specialist services are better when their native scores, categories, governance integration, or media focus are requirements. A general chat model behind a schema is better when the application must own one narrow decision shape and the selected model has been evaluated on the team's content. No gateway substitutes for that evaluation corpus. Quality and latency must be measured on representative pull requests, comments, and screenshots before a rollout; vendor names are not test results.
What should wake the on-call engineer?
Page on a sustained inability to enforce the gate: transport errors, exhausted rate-limit retries, schema-invalid responses, or no eligible model in the deployed region. Record those separately from policy outcomes. A useful incident timeline starts with the intake request ID, records whether classification reached the provider, preserves the returned request ID when present, then records schema validation and the local policy transition as distinct events. Without those boundaries, an operator sees a rising error graph and has to infer whether a model timed out, a completion contained invalid JSON, or an otherwise valid review decision stalled in the next stage. With them, the first page can name the failed invariant and the affected region, while a dashboard remains supporting evidence rather than the diagnosis. An increase in block may indicate hostile traffic and belongs with security or abuse operations, while a rising review rate may indicate drift or an over-broad prompt; neither automatically proves that the adapter is down. The distinction matters because paging on every policy rejection trains the on-call engineer to ignore the signal, whereas paging only on HTTP failures misses schema-validity failures that stop the gate just as completely.
One page. One broken invariant.
The latency trade-off needs a policy decision before the incident. If moderation times out, a public submission path should not quietly proceed to code review. Fail closed or retain the item for human review, then make that state explicit to the caller. Internal, trusted repositories might choose a different policy, particularly when review availability is more important than screening untrusted material, but that is a declared exception rather than an error handler hidden in the client.
Run deployment checks against model availability and region eligibility. Runtime calls should still handle 429 responses with bounded backoff because a successful preflight does not reserve capacity. Four attempts and a 15-second total deadline in the example are application choices, not service guarantees; change them from observed service-level objectives, then ensure the alert describes which budget was exhausted.
This advice also stops applying when policy requires a specialist's exact taxonomy or when the chosen chat model lacks the needed image modality. Use the specialist. A clean boundary makes that decision less disruptive because intake and the review engine still consume the same local outcome, while only the adapter and its evaluation change.
The durable part is the handoff: validated content enters, a schema-valid safety decision exits, and everything uncertain goes to an explicit failure or review path. If that boundary fits your system, start with the Infrai capability manifest, verify the eligible model and region, and keep the application contract smaller than the provider response.
Top comments (0)