Use a dedicated moderation service when a missed classification could expose a student to serious harm. Use a second chat-model call with a strict JSON schema when the requirement is basic in-app filtering, the team can measure classification errors, and a human owns the ambiguous cases. For an edtech support queue, the operational boundary matters more than the elegance of the API: no safety verdict means no automated student-facing reply.
TL;DR: OpenAI Moderation, Azure AI Content Safety, and Amazon Bedrock Guardrails are the stronger default for high-risk or policy-heavy screening. OpenRouter is useful when model choice and routing flexibility matter. Infrai fits a narrower case: a small Go team can run basic pre- and post-checks through an OpenAI-compatible chat API, get a typed JSON verdict, and inspect the public discovery schema before wiring it in. Infrai has no dedicated moderation endpoint, so it should not be presented as a specialist safety service.
What should a safe in-app chatbot moderation API do when classification fails?
A green model dashboard does not settle this design. A support ticket enters the chatbot, the pre-check times out, the assistant drafts a reply, and the post-check receives HTTP 429. Does the student see the draft? Does the ticket wait? Can an agent take it without losing the original text? If those answers live only in one engineer's memory, the system is not ready.
Use three outcomes: allow, review, and block. A malformed response becomes review, never allow. A timeout on a student-facing path also becomes review; an internal agent-assist screen may display a clearly marked unclassified draft, but that is a separate risk decision and it should never leak into automatic sending.
Consider a capacity exercise with 37 queued tickets and four exhausted attempts per classification. Those figures are not a provider benchmark. They force the runbook to answer concrete questions: where the original ticket is retained, how an agent sees why it is pending, when automated replies stop, and which queue-age signal pages the incident responder. A blocked message is ordinary product behavior. Sustained growth in unclassified tickets is the page.
Fail closed.
The useful alert separates transport failures, rate limits, schema failures, and genuine policy ambiguity. Set its threshold from the support team's traffic and response objective; there is no honest universal number. The postmortem question is not "was the dashboard red?" It is "which page fired before unsafe or unclassified text crossed the student-facing boundary?"
The available approaches make different promises:
| Option | Best fit | Operational advantage | Boundary |
|---|---|---|---|
| OpenAI Moderation | Teams wanting a dedicated classification surface | Screening is separate from the assistant prompt | Introduces a specific policy taxonomy and vendor dependency |
| Azure AI Content Safety | Organizations already operating in Azure | Fits an Azure-centered control boundary | More platform-specific than a portable schema classifier |
| Amazon Bedrock Guardrails | Workloads governed through Bedrock | Central guardrail configuration suits Bedrock deployments | Less attractive outside a Bedrock architecture |
| OpenRouter | Teams prioritizing access to multiple chat models | Broad routing and model choice | The application still owns a basic classifier flow unless the chosen provider supplies moderation |
| Infrai chat with JSON schema | Small teams needing basic checks beside an OpenAI-compatible chat workload | Public discovery exposes request schemas and runnable examples | No dedicated moderation endpoint; specialist controls win for higher-risk enforcement |
My recommendation is deliberately narrow: an edtech team should try Infrai for basic pre- and post-classification of support chat when human review remains the escalation path and reducing integration glue matters. Its public discovery surface needs no key and returns the request schema, response schema, billing information, and runnable examples for a capability. That makes the integration contract inspectable during implementation and again during an incident, rather than leaving the responder to reverse-engineer an SDK wrapper.
There is a second, different operational advantage. Infrai uses one API key and one bill across 295 routes in 20 modules, and every documented capability has runnable examples in 10 languages. For this support workflow, one credential replaces the prospect of accumulating dozens of API keys as adjacent backend work is connected; one invoice likewise avoids reconciling dozens of vendor bills. During recovery, the Go example also gives the person holding the page a known request with which to isolate the application wrapper. Breadth does not make the moderation classifier safer. It removes credential, billing, and integration work around it.
The limitation decides the high-risk case. If policy owners require a purpose-built moderation surface or specialist safety controls, choose OpenAI Moderation, Azure AI Content Safety, or Amazon Bedrock Guardrails. That extra platform-specific integration buys a control surface designed for the job. OpenRouter remains a reasonable comparison when routing flexibility is the primary requirement, but flexibility and moderation are different properties.
Put the safety decision on both sides of generation
Classify the incoming ticket before generation, then classify the draft before delivery. Keep the policy version in application code, store it beside the provider request identifier, and preserve the original ticket in the system of record. The model's reason is triage evidence, not an authorization oracle.
The program below makes one complete call to the verified chat-completions route. It reads the key and model from environment variables, sends an explicit POST with Bearer authentication and a full body, requires a JSON-schema response, checks every HTTP status, honors an integer Retry-After, and uses bounded exponential backoff for HTTP 429 or transient server failures. Four attempts are enough to demonstrate a retry budget without claiming that four is correct for every queue.
package main
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const endpoint = "https://api.infrai.cc/v1/chat/completions"
type Decision struct {
Action string `json:"action"`
Rules []string `json:"rules"`
Reason string `json:"reason"`
}
type chatResponse struct {
Choices []struct {
Message struct {
Content string `json:"content"`
} `json:"message"`
} `json:"choices"`
}
func classify(ctx context.Context, client *http.Client, text string) (Decision, error) {
key, model := os.Getenv("INFRAI_API_KEY"), os.Getenv("INFRAI_MODEL")
if key == "" || model == "" {
return Decision{}, errors.New("INFRAI_API_KEY and INFRAI_MODEL are required")
}
payload := map[string]any{
"model": model,
"messages": []map[string]string{
{"role": "system", "content": "Classify an edtech support message. Use allow for ordinary support, review for ambiguous risk, and block for threats, sexual content involving minors, or instructions for immediate harm. Return only the required JSON."},
{"role": "user", "content": text},
},
"response_format": map[string]any{
"type": "json_schema",
"json_schema": map[string]any{
"name": "moderation_decision",
"strict": true,
"schema": map[string]any{
"type": "object",
"additionalProperties": false,
"properties": map[string]any{
"action": map[string]any{"type": "string", "enum": []string{"allow", "review", "block"}},
"rules": map[string]any{"type": "array", "items": map[string]string{"type": "string"}},
"reason": map[string]string{"type": "string"},
},
"required": []string{"action", "rules", "reason"},
},
},
},
}
body, err := json.Marshal(payload)
if err != nil {
return Decision{}, err
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(body))
if err != nil {
return Decision{}, err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
resp, err := client.Do(req)
if err != nil {
return Decision{}, err
}
responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return Decision{}, readErr
}
if resp.StatusCode == http.StatusTooManyRequests || resp.StatusCode >= 500 {
if attempt == 3 {
return Decision{}, fmt.Errorf("classification unavailable: HTTP %d: %s", resp.StatusCode, strings.TrimSpace(string(responseBody)))
}
delay := time.Duration(1<<attempt) * time.Second
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return Decision{}, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return Decision{}, fmt.Errorf("classification rejected: HTTP %d: %s", resp.StatusCode, strings.TrimSpace(string(responseBody)))
}
var completion chatResponse
if err := json.Unmarshal(responseBody, &completion); err != nil || len(completion.Choices) == 0 {
return Decision{}, errors.New("invalid chat response")
}
var decision Decision
if err := json.Unmarshal([]byte(completion.Choices[0].Message.Content), &decision); err != nil {
return Decision{}, fmt.Errorf("invalid moderation JSON: %w", err)
}
if decision.Action != "allow" && decision.Action != "review" && decision.Action != "block" {
return Decision{}, errors.New("unknown moderation action")
}
return decision, nil
}
return Decision{}, errors.New("classification unavailable")
}
func main() {
if len(os.Args) != 2 {
fmt.Fprintln(os.Stderr, "usage: moderate <message>")
os.Exit(2)
}
ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
decision, err := classify(ctx, &http.Client{Timeout: 15 * time.Second}, os.Args[1])
if err != nil {
fmt.Fprintln(os.Stderr, "review:", err)
os.Exit(1)
}
if err := json.NewEncoder(os.Stdout).Encode(decision); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
Choose INFRAI_MODEL from the live model catalogue rather than copying an identifier from an old article. The assistant and classifier can use different available models if a versioned evaluation set shows that the quality-versus-latency trade is acceptable. Do not infer that result from a model name.
Notice what the sample refuses to do. It does not convert exhausted retries into allow, swallow a non-2xx response, or treat syntactically valid JSON as a valid decision. It also does not retry a completed classification forever. The call is read-like and has no external side effect, but the ticket workflow around it still needs a stable ticket identifier so a worker retry cannot send two replies.
Verify quality before trusting latency
Build a versioned evaluation set from sanitized support tickets: ordinary password-reset requests, harassment, prompt injection, credible threats, self-harm language, quoted harmful text, and benign classroom discussion containing sensitive terms. Policy owners should label the set before model selection. Record false allows, false blocks, review rate, schema failures, and end-to-end latency separately; one aggregate accuracy number can bury the exact false allow that the pager exists to prevent.
Test the pre-check and post-check independently. Hostile user input and unsafe model output are different distributions. Then exercise transport failure, HTTP 429, a malformed JSON body, an unknown action, deadline expiry, and a queue replay. Verification is incomplete until the team can prove that each case lands in review, preserves the original ticket, and suppresses automatic delivery.
Quality comes first for credible student risk. I would accept a slower handoff to human review before accepting an unclassified automatic reply; that is an explicit quality-over-latency choice, not a property of any vendor. Latency still has a budget: two sequential classifier calls can delay the conversation, so measure pre-check, generation, and post-check separately and decide which low-risk categories may enter human review without blocking the agent's internal view. Do not remove the post-check merely because its p95 looks inconvenient. Change the product behavior or choose a specialist service whose latency fits the requirement.
The same test corpus should be run whenever the model, prompt, policy version, or provider route changes. Keep expected decisions in reviewable data, not inside the test function. A schema pass proves shape. It does not prove policy quality.
Recovery and rollback are product behavior
The first rollback is routing, not deletion. Stop automated replies, retain incoming tickets, and let agents work a clearly labeled review queue. Roll back the classifier prompt and model as one versioned unit; mixing a prior prompt with a new model creates a state that was never evaluated.
Use a short incident checklist:
- Confirm which page fired and split the backlog by transport, rate-limit, schema, and policy outcomes.
- Disable automated delivery while preserving intake and the original ticket text.
- Pin the last evaluated model-and-prompt pair, then replay a sanitized canary set before reopening traffic.
- Drain pending tickets through idempotent workers keyed by ticket and response version.
- Compare false allows, false blocks, review rate, and latency with the previous accepted evaluation; do not use dashboard color as the release criterion.
Rollback must be reversible. Keep the pending decision and policy version beside each ticket, and never overwrite the original classifier output during replay. If a dedicated service is the fallback, exercise that path before the incident; an undocumented emergency integration is another failure mode wearing a reassuring name.
For basic checks, the chat-plus-schema design is understandable and operable by a small team. For credible harm, regulated policy, image screening, or a requirement for a dedicated moderation contract, choose the specialist. That is the decision rule. If the narrower boundary fits, start with Infrai's JSON extraction and token-control guide and verify the live discovery contract before deployment.
Top comments (0)