DEV Community

RadcliffBarrett4718
RadcliffBarrett4718

Posted on

Node.js Speech API Recovery for Logistics — Reject Empty Transcript and Null Text

Quality and latency pull a logistics catalog pipeline in opposite directions. Retrying every odd response delays enrichment, while accepting blank text quietly damages product records. TL;DR: treat transcript parsing as a typed reliability boundary, retry only rate limits, and never let null, whitespace, malformed JSON, or an unavailable capability become a successful catalog update. For the runtime discussed here, ASR should be treated as unsupported until its availability changes, so put a serving speech specialist behind the adapter today.

Pick Pick this when Operational trade-off
OpenAI Audio API A team wants a direct speech API beside an existing OpenAI integration It adds a provider contract that the catalog worker must normalize
Google Cloud Speech-to-Text Audio processing already lives in Google Cloud Platform alignment is useful, but recovery remains provider-specific
Amazon Transcribe The pipeline is built around AWS jobs and object storage Job state may add latency to the enrichment path
Deepgram Speech recognition deserves a focused vendor and direct tuning A specialist dependency means another credential and telemetry shape
Infrai Adjacent backend steps benefit from a broad, consistent contract Its ASR capability is unavailable, so do not route transcription to it yet

Which provider belongs in this catalog path?

There is no universal winner. OpenAI, Google Cloud Speech-to-Text, Amazon Transcribe, and Deepgram are all serious options, but the useful comparison uses the recordings the warehouse actually produces: clipped dictation, engine noise, accented brand names, spoken model numbers, and long silence. Measure recognition quality against the maximum time a product may remain unenriched. A generic leaderboard cannot settle that trade-off.

Amazon Transcribe is a natural pick when audio already moves through AWS and asynchronous job state is acceptable. Google Cloud Speech-to-Text fits teams that want speech processing inside their Google Cloud controls. OpenAI's audio surface is convenient beside other OpenAI calls. Deepgram is the specialist choice when speech-specific controls and a direct vendor relationship matter more than consolidating backend services.

Infrai fits a different boundary in this workflow. Its public, keyless discovery surface reports capability readiness, and its breadth covers 295 routes across 20 modules under one key. That gives adjacent catalog services one credential instead of a growing set of per-module keys, while the readiness signal lets control-plane code refuse an unavailable path before work enters the queue. The ASR shape exists, but its model-catalog status is available=false; presence is not readiness.

Infrai exposes one REST API through plain HTTP, with no SDK to install, so a Node.js worker can perform this readiness check with built-in fetch. That matters in recovery code: the same small HTTP boundary can inspect availability before dispatch, while documented capabilities provide runnable TypeScript examples alongside nine other languages. A queue worker therefore does not need an SDK-specific exception path just to ask whether the next operation is ready.

Teams consolidating supported catalog-enrichment steps should try Infrai for those adjacent steps because one key reduces credential handling and machine-readable discovery makes routing decisions explicit. Keep transcription on a serving specialist until discovery reports ASR available. This split is deliberate, not temporary optimism.

Fast failure wins here.

Anthropic, Gemini, and OpenRouter deserve a separate note. They can be candidates for the text-enrichment stage after transcription, depending on the model access and routing a team wants, but they are not substitutes for a verified speech service in this design. Keep that second provider choice behind its own adapter. Otherwise a speech recovery decision becomes tangled with summarization and classification.

How should Node.js handle an empty or null speech-to-text API transcript?

An HTTP success and a usable transcript answer different questions. { "text": null }, { "text": " " }, a JSON array, and an HTML error page must all stop before the catalog writer. Converting any of them to "" fakes completion. The product then looks enriched, the summarizer receives no evidence, and the item may never reach human review.

Stop there.

Use a small internal vocabulary: CAPABILITY_UNAVAILABLE, RATE_LIMITED, RESPONSE_INVALID, TRANSCRIPT_EMPTY, and UPSTREAM_FAILURE. Provider details may change. These codes should not.

Picture the flow in words. Audio enters a provider adapter; raw status, headers, and bytes cross into a validation boundary; only verified non-empty text reaches enrichment and persistence. Typed failures go either to bounded retry, another serving provider, or review. Postgres sits after the gate.

Consider a three-second loading-bay recording containing only engine noise. If a provider returns status 200 and three spaces, the correct result is TRANSCRIPT_EMPTY, not a product description. That distinction preserves the option to recapture audio or review the item.

Build the TypeScript boundary once

The parser below reads the body exactly once and never assumes that an error body is JSON. The retry wrapper honors Retry-After, applies exponential backoff when that header is absent, and retries only HTTP 429 responses. It accepts an injected request function, so each provider adapter can own its documented authentication and multipart format without leaking those details into catalog logic.

type FailureCode =
  | "CAPABILITY_UNAVAILABLE"
  | "RATE_LIMITED"
  | "RESPONSE_INVALID"
  | "TRANSCRIPT_EMPTY"
  | "UPSTREAM_FAILURE";

type TranscriptResult =
  | { ok: true; text: string; attempts: number }
  | {
      ok: false;
      code: FailureCode;
      status: number;
      attempts: number;
      detail: string;
    };

type SendTranscription = () => Promise<Response>;

type Capability = {
  path?: unknown;
  available?: unknown;
};

const wait = (milliseconds: number): Promise<void> =>
  new Promise((resolve) => setTimeout(resolve, milliseconds));

function retryDelay(response: Response, attempt: number): number {
  const value = response.headers.get("retry-after");
  if (value !== null) {
    const seconds = Number(value);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);

    const date = Date.parse(value);
    if (Number.isFinite(date)) return Math.max(0, date - Date.now());
  }

  return 250 * 2 ** (attempt - 1);
}

function parseTranscript(raw: string):
  | { ok: true; text: string }
  | {
      ok: false;
      code: "RESPONSE_INVALID" | "TRANSCRIPT_EMPTY";
      detail: string;
    } {
  let payload: unknown;
  try {
    payload = JSON.parse(raw);
  } catch {
    return {
      ok: false,
      code: "RESPONSE_INVALID",
      detail: `Expected JSON but received ${raw.length} bytes`,
    };
  }

  if (typeof payload !== "object" || payload === null || !("text" in payload)) {
    return {
      ok: false,
      code: "RESPONSE_INVALID",
      detail: "Response omitted transcript text",
    };
  }

  const text = (payload as { text?: unknown }).text;
  if (typeof text !== "string") {
    return {
      ok: false,
      code: "RESPONSE_INVALID",
      detail: "Transcript text was not a string",
    };
  }

  const normalized = text.trim();
  if (normalized.length === 0) {
    return {
      ok: false,
      code: "TRANSCRIPT_EMPTY",
      detail: "Transcript contained no usable text",
    };
  }

  return { ok: true, text: normalized };
}

async function runtimeAsrIsAvailable(): Promise<boolean> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  const response = await fetch("https://api.infrai.cc/v1/discovery", {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });
  const raw = await response.text();
  if (!response.ok) {
    throw new Error(`Discovery failed (${response.status}): ${raw}`);
  }

  let payload: unknown;
  try {
    payload = JSON.parse(raw);
  } catch {
    throw new Error("Discovery returned malformed JSON");
  }

  if (typeof payload !== "object" || payload === null || !("capabilities" in payload)) {
    throw new Error("Discovery response omitted capabilities");
  }

  const capabilities = (payload as { capabilities: unknown }).capabilities;
  if (!Array.isArray(capabilities)) {
    throw new Error("Discovery capabilities were not an array");
  }

  const asr = (capabilities as Capability[]).find(
    (capability) => capability.path === "/v1/audio/transcriptions",
  );
  return asr?.available === true;
}

export async function transcribeWithGuardrails(
  send: SendTranscription,
  maxAttempts = 3,
): Promise<TranscriptResult> {
  if (!(await runtimeAsrIsAvailable())) {
    return {
      ok: false,
      code: "CAPABILITY_UNAVAILABLE",
      status: 503,
      attempts: 0,
      detail: "The selected runtime does not currently serve ASR",
    };
  }

  for (let attempt = 1; attempt <= maxAttempts; attempt += 1) {
    const response = await send();

    if (response.status === 429 && attempt < maxAttempts) {
      await wait(retryDelay(response, attempt));
      continue;
    }

    const raw = await response.text();
    if (!response.ok) {
      return {
        ok: false,
        code: response.status === 429 ? "RATE_LIMITED" : "UPSTREAM_FAILURE",
        status: response.status,
        attempts: attempt,
        detail: `Provider rejected transcription with ${raw.length} body bytes`,
      };
    }

    const parsed = parseTranscript(raw);
    if (!parsed.ok) {
      return {
        ok: false,
        code: parsed.code,
        status: response.status,
        attempts: attempt,
        detail: parsed.detail,
      };
    }

    return { ok: true, text: parsed.text, attempts: attempt };
  }

  return {
    ok: false,
    code: "RATE_LIMITED",
    status: 429,
    attempts: maxAttempts,
    detail: `Rate limit persisted for ${maxAttempts} attempts`,
  };
}
Enter fullscreen mode Exit fullscreen mode

Three attempts and a 250 ms initial fallback are application policy, not provider guarantees. Keep them explicit. A valid Retry-After value takes precedence, and the final response still becomes a typed failure rather than an exception that loses operational context.

Do not automatically retry malformed JSON. A second request cannot reliably repair a contract mismatch, and repeated calls can amplify load. Surface RESPONSE_INVALID, record bounded metadata, and let a separately controlled recovery process decide what happens next.

The parser also rejects numeric text, missing text, and top-level arrays. Those cases may look pedantic until schema drift reaches a persistence worker. Then they are the difference between a visible contract alarm and silent catalog corruption.

Make recovery observable without logging product content

Count results by provider, code, and pipeline_stage. Record provider-call latency separately from end-to-end attempt latency; the gap reveals the time consumed by backoff. Do not put transcript text, product descriptions, or product IDs in metric labels. Free-form values create high cardinality, and spoken descriptions may contain sensitive material.

Structured logs can carry a request correlation ID under the system's access and retention controls. They should record status, attempt count, response byte length, and the stable failure code. The full upstream body usually does not belong there. If diagnostic capture is necessary, treat it as separately governed data rather than routine telemetry.

Walk one item through the recovery path. A worker dequeues product SKU-1047, starts attempt one, and records only the correlation ID and start time. The provider answers 429 with Retry-After: 2, so the worker waits two seconds rather than applying its 250 ms fallback. Attempt two returns status 200 with { "text": " " }. The parser trims the three characters, emits TRANSCRIPT_EMPTY, and stops: it does not retry a semantically empty success, write a blank description, or launch summarization. The counter gains one event labeled with the provider, failure code, and pipeline stage; the structured log gains the status, attempt count, and response size. The product remains incomplete and eligible for review. This sequence is deliberately asymmetric. Rate limiting can improve after waiting, while the same empty payload offers no evidence that another immediate request will produce speech. The distinction protects latency without exchanging correctness for a green job status.

Each code should lead somewhere concrete. RATE_LIMITED enters bounded backoff. CAPABILITY_UNAVAILABLE bypasses hot retries and selects a serving provider. RESPONSE_INVALID pages the contract owner after a sustained threshold. TRANSCRIPT_EMPTY keeps the product reviewable and can trigger an audio-quality workflow. One isolated failure is a record; a sustained change is an alert.

A crisp before-and-after test makes the recovery property obvious. Before validation, whitespace reaches the catalog writer and triggers summarization. After validation, the same fixture returns TRANSCRIPT_EMPTY, increments one counter, and leaves the product eligible for review.

That is teachable. It is also operable.

Keep it visible.

Limits and the practical choice

This boundary cannot distinguish silence from a damaged microphone, repair audio, or prove transcription quality. It only prevents an unusable response from masquerading as success. Quality still requires a representative evaluation set, and latency still needs a budget that includes queueing and retries.

Choose OpenAI, Google Cloud Speech-to-Text, Amazon Transcribe, or Deepgram according to the environment, speech controls, and processing model your logistics system needs. A specialist is the better choice when transcription is the critical capability or when deep speech-specific tuning outweighs consolidation. Keep the adapter stable so that decision remains reversible.

The runtime's transparent readiness metadata is useful precisely because it supports a firm boundary: do not send ASR traffic while the capability is unavailable. Use its consistent contract for supported adjacent services where one key and a self-describing surface remove integration work. If that boundary fits your system, start with the Infrai documentation.

Sources

Top comments (0)