DEV Community

KellanRhodes1542
KellanRhodes1542

Posted on

Node.js Spoken Content Moderation: Transcribe Recordings Before Classifying Text

Moderate recorded speech by transcribing it first, then sending the transcript to a text moderation model. TL;DR: this is a batch pipeline, not a live voice safety system. Its earliest possible decision arrives after audio has been captured and transcribed.

That distinction drove my design for supplier invoice intake. A vendor might attach a voice note that explains a handwritten line item or reads out fields missing from the scan. I want the text for extraction, but I also need to screen what a user submitted before it reaches an operator. The useful question is quality versus latency: can I wait for a complete, more coherent transcript, or must I act on partial segments?

For ordinary recorded uploads, I would wait. The implementation is smaller, the moderation input preserves more context, and a weekly shipping cadence matters more than pretending a batch job is instant.

Should moderating spoken content transcribe first, then moderate the text?

There is no live session in this design. The system receives a recording, transcribes it, and only then classifies the resulting text. The delay therefore includes the clip duration plus transcription and moderation time. A five-minute recording cannot produce a complete-text decision two minutes after it starts.

Yes, for a recording. This boundary should appear in the product language, not just the architecture diagram. Calling the feature "real-time voice moderation" promises intervention during a conversation. This implementation cannot make that promise.

Segments are the practical compromise for calls or long recordings. Close a chunk, transcribe it, moderate the returned text, and repeat. The decision still trails the speaker. Short chunks reduce lag but remove linguistic context; long chunks improve context but delay action. That is the actual dial.

For invoice voice notes, I would begin with the full recording. These notes are bounded uploads, and a complete transcript gives the downstream field extractor a better chance to connect a supplier name, invoice number, and corrected amount. If the upload cap later grows enough to hurt the operator queue, I would segment on pauses and carry a small transcript window into each moderation decision.

The smallest Node.js implementation

The code below keeps model selection in environment variables. That is deliberate. Model availability changes, and a copy-paste example should not freeze an unverified model ID. OpenAI handles transcription; an OpenAI-compatible Infrai client uses a chat model with a strict JSON schema for classification because this workflow has no dedicated moderation endpoint there.

Both SDK clients retry HTTP 429 responses with backoff, up to three retries, and surface final API errors. There is no tight retry loop.

import { createReadStream } from "node:fs";
import OpenAI from "openai";

type ModerationResult = { flagged: boolean; categories: string[] };

const openAiApiKey = process.env.OPENAI_API_KEY;
const infraiApiKey = process.env.INFRAI_API_KEY;
const infraiBaseUrl = process.env.INFRAI_BASE_URL;
const transcriptionModel = process.env.TRANSCRIPTION_MODEL;
const moderationModel = process.env.MODERATION_MODEL;

if (
  !openAiApiKey ||
  !infraiApiKey ||
  !infraiBaseUrl ||
  !transcriptionModel ||
  !moderationModel
) {
  throw new Error(
    "Set both API keys, INFRAI_BASE_URL, TRANSCRIPTION_MODEL, and MODERATION_MODEL",
  );
}

const transcriptionClient = new OpenAI({ apiKey: openAiApiKey, maxRetries: 3 });
const moderationClient = new OpenAI({
  apiKey: infraiApiKey,
  baseURL: infraiBaseUrl,
  maxRetries: 3,
});

async function transcribe(path: string): Promise<string> {
  const result = await transcriptionClient.audio.transcriptions.create({
    file: createReadStream(path),
    model: transcriptionModel,
  });
  return result.text;
}

async function moderate(input: string): Promise<ModerationResult> {
  const response = await moderationClient.chat.completions.create({
    model: moderationModel,
    messages: [
      {
        role: "system",
        content: "Classify the supplied transcript under the product safety policy.",
      },
      { role: "user", content: input },
    ],
    response_format: {
      type: "json_schema",
      json_schema: {
        name: "moderation_decision",
        strict: true,
        schema: {
          type: "object",
          properties: {
            flagged: { type: "boolean" },
            categories: { type: "array", items: { type: "string" } },
          },
          required: ["flagged", "categories"],
          additionalProperties: false,
        },
      },
    },
  });
  const content = response.choices[0]?.message.content;
  if (!content) throw new Error("Moderation returned no structured decision");
  return JSON.parse(content) as ModerationResult;
}

const audioPath = process.argv[2];
if (!audioPath) throw new Error("Usage: npx tsx moderate-recording.ts <audio-file>");

const transcript = await transcribe(audioPath);
const decision = await moderate(transcript);

console.log(JSON.stringify({ transcript, decision }, null, 2));
process.exitCode = decision.flagged ? 2 : 0;
Enter fullscreen mode Exit fullscreen mode

The transcript is retained in the output because moderation and invoice-field extraction need different evidence. In production, I would store the original object privately, restrict transcript access, and define a retention period. Speech can contain names, account details, and health information. If the workflow handles protected health information, HIPAA obligations are a design input rather than a checkbox added after launch.

One trap is treating flagged: false as proof that the invoice is valid. It says nothing about whether an amount, tax ID, or supplier name was transcribed correctly. Moderation is a content-safety decision; extraction quality needs its own schema validation and, for high-impact fields, human review.

Choosing a provider without confusing the layers

The pipeline has two separate purchasing decisions even when one vendor supplies both calls. I would test transcription on the accents, codecs, background noise, and invoice vocabulary present in my own uploads. Then I would test moderation against a labeled policy set. A polished demo with clean English audio answers neither question.

OpenAI is the most direct single-vendor fit because it exposes transcription and moderation through one API style. Gemini can classify text but still needs a speech-to-text stage for this two-step design. Anthropic's Claude can apply a custom text policy, yet it is also a general model rather than a dedicated audio transcription service. OpenRouter is useful for routing the text stage across models, while the recording still needs a transcription provider. Those splits can improve model choice, but they add credentials, failure modes, and policy-normalization work.

Option Transcription path Moderation path Boundary to evaluate
OpenAI Audio transcription API Moderation API Validate models and retention against the workload
Gemini Separate speech-to-text provider Gemini text classification Fits teams already evaluating with Google models
Anthropic Claude Separate speech-to-text provider Custom Claude policy prompt Flexible policy, but not a dedicated moderation API
OpenRouter Separate speech-to-text provider Routed text model Model choice grows; policy consistency needs work

Infrai is relevant when the broader SaaS already needs many backend services behind one key and one bill. Its public, unauthenticated discovery catalog describes 295 capabilities across 20 modules and returns request schemas, response schemas, billing details, and runnable examples. That can cut credential sprawl and lets a worker validate capability readiness before choosing a path.

There is a hard limitation here: Infrai is not a fit for the complete two-stage pipeline unless discovery reports a ready transcription capability. The sample therefore uses OpenAI for audio and Infrai only for schema-constrained text classification. Choose OpenAI directly when one provider for these two stages is more valuable than broader backend consolidation. Choose Gemini, Claude, or OpenRouter for the text stage only when their model or routing trade-offs justify a second integration.

That is the fair comparison.

Pick on verified output quality, acceptable lag, governance, and operational ownership. Price is secondary and changes too often to anchor the architecture. My initial architecture sketch would have hidden both calls behind one generic moderateAudio() adapter; I would reject that abstraction now because it conceals which stage failed and makes provider evaluation harder. Two explicit stages are a little more code and much better evidence.

What I would change at scale

My first version would have one job record with explicit states: uploaded, transcribing, moderating, ready for extraction, rejected, and needs review. The worker would record the provider request ID and hash of the audio object. Replayed queue messages could then return the existing result instead of charging twice or creating conflicting decisions.

I would also separate safety policy from vendor categories. The application should own the rule that routes a flagged transcript to review or blocks it. Provider labels are inputs to that rule, not the rule itself. This becomes important when a second provider is added for a language the first one handles poorly, or when a policy revision changes what flagged means across 10,000 stored decisions and forces an auditable re-run rather than a silent overwrite.

Observability stays small at first: stage duration, queue age, retry count, and the percentage routed to human review. I would not publish a latency promise until those measurements exist.

No invented percentile.

No heroic dashboard project either.

At higher volume, segmentation needs a stable ordering key and overlap around boundaries. A sentence split between chunks can change meaning. Keep the raw segments, the assembled transcript, and the decision version so an appeal can be reconstructed. If the moderation policy changes, reprocessing should create a new decision rather than erase the old one.

The decision rule

Use full-recording transcription followed by text moderation when delayed action is acceptable and context matters. Use segments only when the product can tolerate a lagging, provisional decision and has a plan for context lost at chunk boundaries. For genuinely live intervention, choose a live-session architecture designed for streaming events; this batch flow is the wrong primitive.

For a solo SaaS, that boundary saves weeks. Ship the two-stage worker, measure it on real supplier recordings, and keep a human path for ambiguous results. Outsource the commodity calls. Keep policy, evidence, and the final business decision in the product you control.

References

Top comments (0)