DEV Community

evanshepherd5623
evanshepherd5623

Posted on

A Guide to Startup App Moderation Categories for Harassment and PII

A shared gaming platform has one constraint that changes the design: every moderation decision and model call must be attributable to the tenant that submitted it. Short answer: start with seven business categories—harassment, sexual content, self-harm, violence, illegal activity, spam, and privacy/PII exposure—then map each label to allow, review, or block outside the model.

That separation matters when the same platform scores job candidates against a role rubric. Moderation protects candidate notes, chat, and free-text submissions; it must not quietly become another scoring criterion. Keep the hiring rubric, moderation label, enforcement action, and tenant cost record as separate fields.

Keep those apart.

How Should a Startup App Define Moderation Categories for Harassment?

Before: one prompt tries to recognize harmful text, interpret company policy, reject the submission, and explain the decision. A policy change means prompt surgery. Audit records often preserve only the final verdict.

After: the model emits a small, stable label plus evidence and confidence. Application code applies the tenant's policy matrix. The scoring worker receives content only after that gate, while billing metadata is joined to the same tenantId and requestId.

Picture the flow in words: candidate text enters, classification produces facts, tenant policy chooses an action, and only allowed content reaches rubric scoring. Review items stop in a human queue. The cost event travels beside the decision, not inside it.

Seven labels are enough for a first release. More labels can feel precise, but every extra boundary creates another prompt distinction and another reviewer choice. Expand only when a real policy needs a different action. For example, splitting PII into contact details and government identifiers is useful only if the product handles them differently. The concrete trade-off is recall against reviewer clarity: a broad pii label loses detail, while five PII subtypes make the queue slower to scan and the prompt easier to misread. Begin with the broad label, examine reviewer overrides, and split it only when two subtypes need different handling.

The action table belongs in ordinary application configuration:

Category Typical starting action Why it stays configurable
Harassment Review Context and target matter
Sexual content Review or block Audience and product policy differ
Self-harm Review Escalation policy is product-specific
Violence Review or block Discussion is not the same as threat
Illegal activity Review Jurisdiction and intent matter
Spam Block Repetition and solicitation rules vary
Privacy/PII Block Exposure rules depend on data handling

These are starting actions, not universal moral judgments. A category is an observation. An action is a business decision.

A small contract beats a clever prompt

The following TypeScript example calls an OpenAI-compatible chat endpoint with JSON Schema output. It uses one write route, an explicit method, status checks, and bounded retry behavior for 429. Set AI_BASE_URL to the provider's versioned API base, obtain AI_MODEL from that provider's current model catalog, and keep the key in the environment.

type Category =
  | "harassment"
  | "sexual"
  | "self_harm"
  | "violence"
  | "illegal"
  | "spam"
  | "pii"
  | "none";

type Classification = {
  category: Category;
  confidence: number;
  evidence: string;
};

const schema = {
  name: "moderation_classification",
  strict: true,
  schema: {
    type: "object",
    additionalProperties: false,
    required: ["category", "confidence", "evidence"],
    properties: {
      category: {
        type: "string",
        enum: [
          "harassment",
          "sexual",
          "self_harm",
          "violence",
          "illegal",
          "spam",
          "pii",
          "none",
        ],
      },
      confidence: { type: "number", minimum: 0, maximum: 1 },
      evidence: { type: "string", maxLength: 240 },
    },
  },
} as const;

const requiredEnv = (name: string): string => {
  const value = process.env[name];
  if (!value) throw new Error(`Missing ${name}`);
  return value;
};

const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

async function classify(text: string): Promise<Classification> {
  const baseUrl = requiredEnv("AI_BASE_URL").replace(/\/$/, "");
  const apiKey = requiredEnv("INFRAI_API_KEY");
  const model = requiredEnv("AI_MODEL");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${baseUrl}/chat/completions`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({
        model,
        messages: [
          {
            role: "system",
            content:
              "Classify the text into exactly one supplied category. Report a short verbatim evidence span. Do not choose an enforcement action.",
          },
          { role: "user", content: text },
        ],
        response_format: { type: "json_schema", json_schema: schema },
      }),
    });

    if (response.status === 429 && attempt < 3) {
      const retryAfter = Number(response.headers.get("retry-after"));
      const delayMs = Number.isFinite(retryAfter)
        ? retryAfter * 1_000
        : 500 * 2 ** attempt;
      await sleep(delayMs);
      continue;
    }

    if (!response.ok) {
      throw new Error(`Classification failed (${response.status}): ${await response.text()}`);
    }

    const payload = (await response.json()) as {
      choices: Array<{ message: { content: string } }>;
    };
    return JSON.parse(payload.choices[0].message.content) as Classification;
  }

  throw new Error("Classification remained rate limited after four attempts");
}

const policy: Record<Category, "allow" | "review" | "block"> = {
  harassment: "review",
  sexual: "block",
  self_harm: "review",
  violence: "review",
  illegal: "review",
  spam: "block",
  pii: "block",
  none: "allow",
};

const result = await classify("Candidate submission text goes here");
console.log({ classification: result, action: policy[result.category] });
Enter fullscreen mode Exit fullscreen mode

Store tenantId, requestId, taxonomy version, category, confidence, action, and reviewer override together. Do not put the tenant name or candidate identity into the category enum. The schema should remain boring.

The contract has 8 enum values, a confidence range from 0 to 1, and a 240-character evidence cap. The retry loop makes no more than 4 attempts. I would keep those limits visible in code review because each one changes an operational outcome: queue shape, thresholding, log volume, or failure time.

Infrai combines a plain REST API, one key, and one bill across 295 routes in 20 modules; its specified per-call cost, vendor, latency, and request metadata can support tenant attribution without installing another SDK. That unified account reduces credential and invoice reconciliation work when each gaming tenant uses more than moderation. The genuinely self-describing API exposes request and response schemas through public discovery without requiring a key. It does not provide a dedicated moderation endpoint, so the chat-model-plus-JSON-Schema pattern is the relevant path. That is a real tradeoff: the contract is portable, but policy quality and evaluation remain your responsibility.

Which provider shape fits the workflow?

There is no honest winner without knowing how much policy control and cloud alignment the team needs. Compare interfaces before brands.

Option Interface shape Best fit Boundary to plan for
OpenAI Moderation Dedicated moderation model and endpoint Teams that want provider-defined safety categories Translate provider labels into your business taxonomy
Azure AI Content Safety Dedicated content-safety service Teams already operating inside Azure governance Keep Azure severity results separate from hiring scores
Amazon Comprehend Managed text analysis APIs AWS-centered pipelines that want managed text classification tools A business moderation taxonomy still needs an explicit mapping layer
Anthropic Claude General chat model with structured application output Teams already using Claude for adjacent text workflows Own and evaluate the moderation taxonomy in the application
Google Gemini General model available in Google's AI stack Teams consolidating model work around Google Keep model classification separate from enforcement policy
OpenRouter Multi-model routing interface Teams comparing model behavior behind one integration Provider variation makes schema validation and evaluation important
Infrai Chat model with JSON Schema through an OpenAI-compatible REST surface Teams prioritizing one HTTP contract and per-call attribution No dedicated moderation endpoint; validate the schema and policy yourself

Google Cloud Natural Language is another real text-analysis product worth evaluating for a Google Cloud estate, but a generic text-analysis result is not automatically a moderation decision. Amazon Rekognition Moderation Labels is relevant when candidate uploads include images rather than only text. Modality changes the comparison.

Choose the dedicated service when its native labels match the policy closely. Choose the structured chat approach when the business taxonomy must stay small, explicit, and portable. In either case, run a labeled evaluation set from the actual gaming community before allowing the result to affect production flow. No invented benchmark can replace that test.

What about ambiguous or multi-label content?

A single-label contract is intentionally restrictive. It makes routing and reviewer queues easy to reason about, but text can contain both PII and harassment. Pick the category with the stricter configured action for the first pass, preserve the evidence, and let the reviewer add secondary labels. If secondary labels repeatedly change outcomes, that is evidence for a schema revision.

Do not route low confidence straight to allow. Send it to review. Confidence is a model output, not a calibrated guarantee, so choose the threshold from a labeled evaluation set and record the taxonomy version used in that experiment.

There is another edge: quoted or educational discussion. A candidate might analyze violent game dialogue without threatening anyone. Evidence spans help a reviewer see why the classifier fired, while the separate action layer lets tenant policy retain context. Keep the raw candidate content under the product's existing access and retention rules; avoid duplicating it in cost logs.

Can moderation bias the hiring score?

Yes, if the pipeline lets safety labels leak into the job rubric. Prevent that by making moderation a content-handling gate, not a feature supplied to the scoring model. A blocked submission does not become a low-scoring candidate. It becomes a policy event requiring the configured handling path.

This boundary also sharpens observability. The moderation dashboard should answer: which tenant generated the call, which taxonomy version ran, what action followed, and how often reviewers overrode it? The hiring dashboard should answer whether rubric dimensions were applied consistently. Joining them for an authorized investigation is useful; blending their metrics by default is risky.

Per-tenant cost visibility follows the same rule. Capture provider cost metadata where available, attach it to the tenant and request, then aggregate outside candidate evaluation. Cost can choose a routing policy. It should never change a candidate's score.

Start with 7 labels, 3 actions, and 1 frozen taxonomy version.

Test it on representative content, inspect disagreements, and add a category only when it changes a real operational decision. Small is maintainable.

Sources

Top comments (0)