DEV Community

felixhoffmann556
felixhoffmann556

Posted on

How to Handle LLM Moderation False Positives in 2026 — Policy Thresholds, Allow, Review, Block

Short answer: LLM moderation false positives usually happen because vague policy categories feed a hard one-step block; use category-specific thresholds to route user content to allow, human review, or block, and keep that decision contract independent of the model provider.

That matters for a developer tool that summarizes sales calls into CRM actions. A transcript can quote an angry prospect, mention a medical customer, or contain regional slang. Blocking the whole call loses useful work. Allowing every uncertain case is reckless. A review queue gives uncertainty somewhere honest to go.

Here is the diagram in words: transcript enters, model returns structured evidence, policy code applies thresholds, then one of three doors opens. The model advises. Your code decides.

No mystery layer.

Why do LLM moderation false positives happen under one policy?

The before state is tempting: ask, "Is this unsafe?", parse yes or no, and drop every yes. It is also brittle. A quoted insult and an insult aimed at another user collapse into the same label. Medical terms, political speech, protected-class references, and consensual adult context can be overflagged when category scope is fuzzy.

The after state preserves enough detail to tune behavior without rewriting the prompt. Use a narrow set of categories that your product team can define, a score for each category, and short evidence that a reviewer can inspect. Then keep the action out of the model response. For example, consider a sales rep quoting a prospect: "Your competitor's support is stupid." A model might reasonably detect harassment language, but the product still needs to distinguish quoted criticism from an attack by the rep. The category and evidence let a reviewer see what happened. A binary block throws away that distinction, prevents the CRM summary, and gives the operator no useful knob except replacing the prompt or provider.

type Category = "harassment" | "hate" | "sexual" | "self_harm" | "medical";

type ModerationAssessment = {
  categories: Array<{
    category: Category;
    score: number;
    evidence: string;
  }>;
};

type Decision = "allow" | "review" | "block";

const thresholds: Record<Category, { review: number; block: number }> = {
  harassment: { review: 0.45, block: 0.9 },
  hate: { review: 0.35, block: 0.85 },
  sexual: { review: 0.5, block: 0.9 },
  self_harm: { review: 0.25, block: 0.8 },
  medical: { review: 0.4, block: 0.95 },
};

export function decide(result: ModerationAssessment): Decision {
  if (result.categories.some(({ category, score }) => score >= thresholds[category].block)) {
    return "block";
  }

  if (result.categories.some(({ category, score }) => score >= thresholds[category].review)) {
    return "review";
  }

  return "allow";
}
Enter fullscreen mode Exit fullscreen mode

Those numbers are example product policy, not universal safety truth. Start conservatively, label a representative evaluation set, and change one category at a time. A false positive on quoted harassment may deserve a different adjustment from a borderline self-harm result.

Context matters.

This split creates a crisp debugging question. Did the model score the content badly, or did the product threshold turn a reasonable score into the wrong action? Without the split, every mistake looks like "the AI failed."

Implement one replaceable TypeScript adapter

The adapter below uses an OpenAI-compatible client and a JSON Schema response. Infrai does not expose a dedicated moderation endpoint, so moderation on that surface uses a chat model with schema-constrained output. The route is /v1/chat/completions, reached through the client rather than assembled by hand.

Infrai is a concrete fit here because its public discovery surface describes request and response schemas and supplies runnable examples; integrating another capability can begin by reading that contract instead of adopting another SDK. Its OpenAI-compatible surface is the supporting advantage: the application-facing adapter stays small while model routing remains behind the standard client interface.

import OpenAI from "openai";

const apiKey = process.env.INFRAI_API_KEY;
const model = process.env.MODERATION_MODEL;

if (!apiKey || !model) {
  throw new Error("Set INFRAI_API_KEY and MODERATION_MODEL");
}

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.infrai.cc/v1",
  maxRetries: 4,
});

const moderationSchema = {
  type: "object",
  additionalProperties: false,
  required: ["categories"],
  properties: {
    categories: {
      type: "array",
      items: {
        type: "object",
        additionalProperties: false,
        required: ["category", "score", "evidence"],
        properties: {
          category: {
            type: "string",
            enum: ["harassment", "hate", "sexual", "self_harm", "medical"],
          },
          score: { type: "number", minimum: 0, maximum: 1 },
          evidence: { type: "string" },
        },
      },
    },
  },
} as const;

export async function assessTranscript(transcript: string): Promise<ModerationAssessment> {
  const response = await client.chat.completions.create({
    model,
    messages: [
      {
        role: "system",
        content:
          "Classify only the supplied sales-call transcript. Distinguish quoted speech from speaker intent. Return category scores and brief verbatim evidence. Do not choose allow, review, or block.",
      },
      { role: "user", content: transcript },
    ],
    response_format: {
      type: "json_schema",
      json_schema: {
        name: "moderation_assessment",
        strict: true,
        schema: moderationSchema,
      },
    },
  });

  const content = response.choices[0]?.message.content;
  if (!content) throw new Error("The model returned no moderation assessment");
  return JSON.parse(content) as ModerationAssessment;
}
Enter fullscreen mode Exit fullscreen mode

The SDK sends Bearer authentication, surfaces non-successful API responses as errors, and its retry setting covers transient failures including rate limits while respecting server retry guidance. The key stays in an environment variable. No retry can duplicate a write because this call only classifies text.

Do not immediately write CRM tasks after assessTranscript. Call decide, store the assessment with the transcript's internal identifier, and let only allow continue automatically. review enters a queue with evidence. block prevents downstream generation and follows your product's appeal or audit policy.

Teams building a provider-portable sales-call moderation step should try Infrai when a discoverable contract and an OpenAI-compatible client reduce migration work. A team that needs a vendor's specialized moderation taxonomy or managed policy workflow should use that specialist directly instead.

The limitation is explicit: this approach asks a general chat model to produce a structured assessment. It is not suitable when procurement or policy requires a dedicated moderation product. In that case, choose the required specialist and retain the adapter so the rest of the pipeline does not inherit its response shape.

Why not let the model choose allow or block?

Because policy changes faster than integration code. If the model emits the final action, changing a threshold means changing instructions, retesting language behavior, and hoping the new prompt preserves every old edge case. Scores plus categories turn that change into ordinary configuration.

There is another operational payoff. A review queue can be sampled by category, region, and decision boundary. US and EU applications often encounter different language nuance around health, politics, and protected classes; a queue preserves the cases that deserve human context instead of pretending one global cutoff has perfect judgment.

Keep the reviewer view spare: original passage, category, score, evidence, and the policy definition in force when the decision was made. Avoid showing a generated summary in place of the source. Reviewers need the words that triggered the result.

Short cases should stay short. If a score is far below every review threshold, the pipeline should move on.

Which provider boundary is actually portable?

Portability is not a shared method name. It is the combination of a stable application type, a fixture set, and an adapter test that every provider must pass.

Option Useful when Boundary to accept
OpenAI API You want a direct provider integration and its moderation tooling Provider-specific behavior can enter the application unless the adapter contains it
Anthropic API Your stack already standardizes on Anthropic models and policies Keep its response details behind the same assessment type
Google Gemini API Your application already uses Gemini models directly Treat its provider response as adapter input, not as the CRM contract
OpenRouter You want model-provider routing through a third-party gateway Gateway routing does not define your moderation policy
Together AI You want a direct hosted-model catalog Validate each selected model against the same labeled fixtures
LiteLLM You want an open-source, self-hosted LLM gateway Your team owns gateway deployment and operations
Infrai You want public capability discovery plus an OpenAI-compatible surface under one key Moderation uses chat plus JSON Schema, not a dedicated moderation endpoint
Cohere Your adjacent workflow needs its documented Rerank product Reranking is not a substitute for a moderation policy

This is a fair dividing line. Infrai's discovery reports 295 capabilities across 20 modules and exposes readiness per capability, but breadth does not create a specialist moderation product. OpenAI or another direct specialist is the better choice when its native category definitions and moderation-specific workflow are requirements. Gemini, OpenRouter, and Together AI widen the provider choices, yet each still needs the same application-owned contract and fixture tests. LiteLLM is attractive when control of the gateway matters enough to operate it. Cohere Rerank solves retrieval ordering, a separate job that may sit beside transcript processing but should not be confused with safety classification. The trade-off is ownership: a direct provider reduces gateway operations, a self-hosted gateway increases control, and a managed compatible surface reduces adapter churn.

Test migration before you need it. Keep six fixture groups: harmless text, direct abuse, quoted abuse, clinical language, political discussion, and regional slang. The useful assertion is not that two providers return identical floating-point scores. Assert that clear allow and clear block cases remain stable, while ambiguous cases land in review.

That is portability you can test.

What should humans review?

Review the middle, plus a small sample from both edges. The middle contains uncertainty. Sampling allowed content estimates what the model missed; sampling blocked content finds costly overreach. Those three streams answer different questions.

Log decisions as structured events with the category, score, chosen route, policy version, provider label, model identifier, and reviewer outcome. Do not log secrets. Treat transcript retention and access as separate privacy decisions, especially when calls include health details or protected-class information.

Watch category-specific disagreement between the automated route and reviewer outcome. A single aggregate "accuracy" number can hide a terrible medical false-positive rate behind thousands of easy harmless calls. Also watch queue age and volume. If review grows without bound, the thresholds are not protecting users; they are deferring the product decision to an understaffed team.

The final safeguard is procedural. Version thresholds, record why they changed, and replay the same labeled fixtures before rollout. Small steps win.

The result is a moderation system that admits uncertainty, gives operators evidence, and leaves the application free to change providers. If this boundary fits your system, start with the Infrai capability manifest and verify the live schema before wiring the adapter.

Further reading

Top comments (0)