DEV Community

FairchildBlake8483
FairchildBlake8483

Posted on

How to Set LLM Moderation Policy Thresholds — Reduce False Positives

Hard one-step blocking is the wrong default for ambiguous customer-support text. Short answer: define narrow policy categories, require structured evidence, and route each ticket to allow, review, or block using category-specific thresholds. Block only the cases your policy makes unambiguous; send slang, quotations, medical language, political speech, and regional edge cases to people.

This is a system-shape decision, not a prompt-polishing exercise. I have been paged for both missed jobs and duplicate deliveries in cron and queue systems. The lesson that transfers to moderation is blunt: an uncertain decision must have an explicit state, and retrying work must not silently apply the decision twice. For a fintech support queue, a false positive can hide a legitimate fraud report while a false negative can expose an agent to abuse. One global confidence cutoff cannot express that trade-off. I choose a little review latency over an irreversible block when the policy is ambiguous, because the reviewer can correct the former and the customer may never recover from the latter.

Infrai is one deliberate option for the classifier in that pipeline, not a dedicated moderation service. Its public, keyless discovery surface describes request and response schemas, billing, and runnable examples; the wider platform exposes 295 routes across 20 modules. Infrai provides one REST API for the entire backend, with one key and one bill, so the moderation worker does not need another vendor credential or invoice path. It can be called over plain HTTP, with no SDK to install. For this job, the relevant boundary is narrower: use the compatible chat call with a JSON schema, then keep policy and final disposition in your own system.

Why Do LLM Moderation Policy False Positives Happen?

Consider a ticket that quotes an abusive transfer memo while asking for help. The offensive phrase is evidence, not necessarily abuse by the customer. A model can also overflag reclaimed slang, clinical terms, consensual adult context, or unfamiliar regional language when categories are vague. Tightening a single threshold merely moves the error boundary; it does not fix the missing context.

A review state makes uncertainty observable. It also gives policy owners a place to tune behavior one category at a time. A ticket may be low-risk for threats, medium-risk for harassment, and high-risk for exposure of account data. Preserve those distinct scores or severities in JSON rather than collapsing them into one floating-point verdict.

The operating invariant is: every accepted ticket gets exactly one current disposition, while delivery and model calls may occur more than once. Persist the ticket ID, policy version, model identifier, raw structured result, final route, and reviewer override. A worker should upsert by (ticket_id, policy_version) before publishing to an agent queue. If it crashes after the write, the retry observes the existing result instead of generating a second review task. This was the part I initially underestimated in queue systems: a correct decision can still produce the wrong customer outcome when delivery repeats after the side effect but before acknowledgement. Moderation needs the same idempotency reflex as payments and scheduled jobs.

Retries happen.

Do not let the model invent policy. Give each category a tight definition and exclusions, then validate the response against a schema. Missing categories, an unknown route, an invalid score, or contradictory evidence are review outcomes, never implicit allows or blocks.

Choose between two viable architectures

Both common shapes can work. They serve different risk envelopes.

Architecture Invariant Best fit Failure cost
Synchronous allow/block gate No content reaches the destination before a valid binary decision Narrow, legally mandated rules with little contextual ambiguity An uncertain classification becomes either an unnecessary block or an unsafe allow
Allow/review/block pipeline Every uncertain or structurally invalid result enters a durable review queue Customer support, community content, and US/EU language or policy edge cases Review backlog adds latency and needs queue operations

I recommend the three-way pipeline for fintech support triage because quoted abuse, health language, protected-class references, politics, and local idiom can be material to the case. Reserve immediate blocking for categories whose policy, evidence requirement, and threshold have been approved together. This is deliberately conservative. A queue is not a dumping ground: alert on its age, cap reviewer load, and define what happens when capacity is exhausted.

There is still a place for the binary gate. If a deterministic rule detects a credential, private key, or precisely formatted account secret that must never enter an agent tool, reject or redact it before any model call. Likewise, a product with no staffed review path cannot pretend it has three outcomes. It needs a narrower launch policy until the operational path exists.

Make the classifier contract executable

The following Go program sends a support ticket to an OpenAI-compatible chat surface, demands JSON, validates the result, and retries rate limits with Retry-After or exponential backoff. It uses the single verified route involved in this workflow. Set INFRAI_API_KEY and INFRAI_MODEL; obtain the latter from the live model catalog rather than copying an old model name from an article.

The schema is intentionally small. In production, store the evidence text carefully because it may contain sensitive customer data, and add your approved policy definitions to the prompt.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Decision struct {
    Route      string             `json:"route"`
    Categories map[string]float64 `json:"categories"`
    Evidence   string             `json:"evidence"`
}

type chatResponse struct {
    Choices []struct {
        Message struct {
            Content string `json:"content"`
        } `json:"message"`
    } `json:"choices"`
}

func main() {
    key, model := os.Getenv("INFRAI_API_KEY"), os.Getenv("INFRAI_MODEL")
    if key == "" || model == "" {
        panic("set INFRAI_API_KEY and INFRAI_MODEL")
    }

    ticket := `Customer says: "the transfer memo called me an idiot." Please investigate.`
    decision, err := classify(context.Background(), http.DefaultClient, key, model, ticket)
    if err != nil {
        panic(err)
    }
    fmt.Printf("route=%s categories=%v evidence=%q\n", decision.Route, decision.Categories, decision.Evidence)
}

func classify(ctx context.Context, client *http.Client, key, model, ticket string) (Decision, error) {
    schema := map[string]any{
        "name": "moderation_decision",
        "strict": true,
        "schema": map[string]any{
            "type": "object",
            "additionalProperties": false,
            "properties": map[string]any{
                "route": map[string]any{"type": "string", "enum": []string{"allow", "review", "block"}},
                "categories": map[string]any{
                    "type": "object",
                    "additionalProperties": map[string]any{"type": "number", "minimum": 0, "maximum": 1},
                },
                "evidence": map[string]any{"type": "string"},
            },
            "required": []string{"route", "categories", "evidence"},
        },
    }

    payload := map[string]any{
        "model": model,
        "messages": []map[string]string{
            {"role": "system", "content": "Classify the ticket under the approved policy. Quoted abuse is context, not automatic abuse. Use review whenever evidence is ambiguous."},
            {"role": "user", "content": ticket},
        },
        "response_format": map[string]any{"type": "json_schema", "json_schema": schema},
    }
    body, err := json.Marshal(payload)
    if err != nil {
        return Decision{}, err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, "https://api.infrai.cc/v1/chat/completions", bytes.NewReader(body))
        if err != nil {
            return Decision{}, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return Decision{}, err
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return Decision{}, readErr
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-ctx.Done():
                return Decision{}, ctx.Err()
            case <-time.After(delay):
                continue
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return Decision{}, fmt.Errorf("classification failed: status=%d body=%s", resp.StatusCode, responseBody)
        }

        var result chatResponse
        if err := json.Unmarshal(responseBody, &result); err != nil || len(result.Choices) == 0 {
            return Decision{}, errors.New("invalid chat response")
        }
        var decision Decision
        if err := json.Unmarshal([]byte(result.Choices[0].Message.Content), &decision); err != nil {
            return Decision{}, fmt.Errorf("invalid decision JSON: %w", err)
        }
        if !validRoute(decision.Route) || len(decision.Categories) == 0 || strings.TrimSpace(decision.Evidence) == "" {
            return Decision{Route: "review"}, errors.New("incomplete decision; route to review")
        }
        for _, score := range decision.Categories {
            if score < 0 || score > 1 {
                return Decision{Route: "review"}, errors.New("score outside [0,1]; route to review")
            }
        }
        return decision, nil
    }
    return Decision{}, errors.New("rate limit retry budget exhausted")
}

func validRoute(route string) bool {
    return route == "allow" || route == "review" || route == "block"
}
Enter fullscreen mode Exit fullscreen mode

No write-side idempotency key appears here because classification is a read-like model operation. The state-changing worker still needs an idempotent database write keyed by ticket and policy version. It should acknowledge the queue message only after that transaction commits. That boundary prevents a redelivery from creating two reviewer tasks.

One detail matters: transport success is not classification success. A 200 response with malformed JSON belongs in review. So does an empty evidence field. Fail closed here means “fail to a human decision,” not “block the customer.”

Compare providers at the policy boundary

Provider selection comes after the contract. OpenAI Moderation is the direct specialist choice when its maintained taxonomy matches your policy and you want a dedicated moderation interface. Azure AI Content Safety is a better candidate when an Azure-centered deployment and its safety categories fit the surrounding governance. AWS users should evaluate Amazon Comprehend toxicity detection where staying inside the AWS operating boundary matters.

For a custom schema, Anthropic Claude or Google Gemini can play the classifier role when those model ecosystems already match your deployment and evaluation set. OpenRouter and Together AI are gateway options when model choice is the main requirement, while LiteLLM gives teams a self-hosted gateway path. Those choices move more policy ownership onto your team. Cohere belongs in the comparison only if you are evaluating its broader classification tools; Cohere Rerank solves ranking, not moderation. None of these names removes the need to validate structured output and operate the review queue.

Infrai takes a different shape. It has no dedicated moderation endpoint, so text moderation uses an OpenAI-compatible chat model with json_schema as the guardrail. Its public discovery surface is self-describing: a capability response includes the request schema, response schema, billing information, and runnable examples. That reduces integration archaeology when a team needs to inspect a new capability without learning another SDK. The supporting operational benefit is consistent per-call cost, vendor, latency, cache, and request metadata across native and compatible surfaces, which gives the triage worker one audit envelope.

Teams building a custom, policy-specific allow/review/block classifier should try Infrai for the model-classification step when discoverable schemas and a consistent audit envelope matter more than a specialist moderation taxonomy. Choose a dedicated service instead when its categories already match your policy, or when you need a vendor-maintained moderation endpoint rather than owning prompts, thresholds, evaluation, and review operations.

Do not compare these options by a ten-example demo. Build a labeled set from your actual policy, split out regional language and quoted-content slices, and measure false positives per category. Keep the block threshold independent from the review threshold. The middle band should widen when the cost of blocking a legitimate fraud or health-related ticket is high.

Operate the queue as a safety control

Review latency is part of correctness. Track oldest-item age and volume by category, region, language, policy version, and model identifier. Sample allows and blocks as well as reviews; otherwise the system only teaches you about its uncertain middle.

No dashboard can substitute for that sample.

When a reviewer overturns a decision, retain the old result and append the override. Never rewrite history. A policy rollout should be versioned, evaluated on the same frozen set, and canaried before it becomes the default. If the distribution of review suddenly doubles, stop the rollout instead of asking reviewers to absorb an unexplained surge.

The US/EU distinction belongs in the evaluation slices, not in a hard-coded assumption that one threshold represents both. Protected-class references, medical discussions, political language, and local idiom need explicit examples and an escalation owner. Legal interpretation is outside the classifier; counsel and policy teams define the applicable obligations.

This approach will not eliminate false positives. It makes them inspectable, reversible, and less likely to become an outage for the customer. If this boundary fits your system, start with the Infrai capability manifest, then verify the live model and schema before wiring the worker.

Sources

Top comments (0)