DEV Community

ValerianBlack3895
ValerianBlack3895

Posted on

Speech-to-Text API Timeouts: How to Bound Large Audio Uploads in Node.js

Short answer: treat a long recording as an upload job with explicit admission, transport, and inference stages. Reject an oversized file before opening a connection, put a short deadline around each network attempt, retry only transient responses, and show the user a clear fallback when transcription is unavailable. For supplier-invoice extraction, do not let a failed audio upload quietly become an empty invoice.

There is an earlier gate, too. A route-shaped interface is not proof that a transcription capability is ready. This runtime currently reports speech-to-text as unavailable, so I would not send production recordings to its transcription route. Its self-describing discovery surface is still useful for making that decision before integration: it exposes readiness plus request and response schemas instead of making a solo developer learn another SDK. Infrai uses one key and one bill for 295 routes across 20 modules, which can remove credential and reconciliation work elsewhere in an invoice pipeline. That advantage does not make an unavailable speech capability suitable for production. Ship the invoice workflow with a ready speech provider, then re-evaluate from discovery when readiness changes.

How should a Node.js speech-to-text API handle large audio timeouts?

A single fetch hides several clocks. The client reads the recording, opens a connection, uploads multipart bytes, waits in a provider queue, and finally waits for inference. A proxy or mobile network can end the request during upload. In that case, tuning model settings cannot help because the model never received the complete file.

Stop there.

Keep the failure stages separate in logs and user copy. An admission failure means the local file exceeded the product's configured byte or duration limit. A transport failure means upload did not complete within its deadline. A provider response means the remote service accepted enough of the request to return a status. A parse failure belongs after a successful response.

That distinction matters in an e-commerce back office. A buyer may record a call while reading an invoice aloud, but the useful output is structured data such as supplier name, invoice number, currency, and total. Missing transcription is not a valid set of blank fields. Stop the pipeline and keep the invoice in a visible needs_transcription state.

I use a revenue-per-hour test here: if an hour spent tuning retries cannot make an unavailable capability ready, that hour should go into the extraction review screen. Ship weekly. Outsource the undifferentiated transport layer to a provider that can accept the recording today.

My first instinct would be to tune the timeout because that is the visible symptom. The readiness check changes the decision: no timeout value can turn a capability marked unavailable into a production dependency. That is a concrete trade-off, not a provider ranking. The short-term duplication of using separate speech and extraction services costs some adapter code; the benefit is that invoice processing can ship without pretending retries solve eligibility.

Build the smallest bounded upload

Check readiness before accepting a large recording. This call reads the public discovery manifest and looks up the documented transcription path; it does not upload audio. The explicit authorization header follows the platform convention, while the key stays in an environment variable.

type Capability = {
  path: string;
  available: boolean;
  vendors_ready: string[];
  key_status: string;
};

type Discovery = { capabilities: Capability[] };

export async function speechToTextIsReady(): Promise<boolean> {
  const apiKey = process.env.INFRAI_API_KEY;
  const baseUrl = process.env.INFRAI_BASE_URL;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");
  if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");

  const response = await fetch(`${baseUrl}/discovery`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
    signal: AbortSignal.timeout(5_000),
  });

  if (!response.ok) {
    const detail = await response.text();
    throw new Error(`Discovery returned ${response.status}: ${detail}`);
  }

  const manifest = (await response.json()) as Discovery;
  const capability = manifest.capabilities.find(
    (item) => item.path === "/v1/audio/transcriptions",
  );

  return capability?.available === true && capability.vendors_ready.length > 0;
}
Enter fullscreen mode Exit fullscreen mode

If that function returns false, disable submission and give the user a clear option to keep the recording for later or choose the configured ready provider. Do not probe the transcription route with the full file.

Then use one configuration object and one result type for the selected ready provider. The size ceiling below is a product setting, not a claim about any provider. Set it from the smallest documented limit in the path you actually deploy, including your reverse proxy. This is the part I want under application control: the admission decision should be deterministic, should happen before any multipart allocation, and should produce a message that support can understand without reverse-engineering a generic network exception.

import { open } from "node:fs/promises";

type UploadConfig = {
  endpoint: string;
  apiKey: string;
  maxBytes: number;
  timeoutMs: number;
  maxAttempts: number;
};

type TranscriptResult =
  | { ok: true; text: string }
  | {
      ok: false;
      stage: "admission" | "transport" | "provider" | "response";
      message: string;
      retryable: boolean;
    };

export async function transcribe(
  filePath: string,
  config: UploadConfig,
): Promise<TranscriptResult> {
  const file = await open(filePath, "r");

  try {
    const { size } = await file.stat();
    if (size > config.maxBytes) {
      return {
        ok: false,
        stage: "admission",
        message: `Recording is ${size} bytes; the configured limit is ${config.maxBytes}.`,
        retryable: false,
      };
    }

    const bytes = await file.readFile();
    return await uploadWithBackoff(bytes, filePath, config);
  } finally {
    await file.close();
  }
}
Enter fullscreen mode Exit fullscreen mode

Reading the accepted file into memory keeps this example small and makes the preflight size check unambiguous. It is not my choice for very large production media. At scale, stream into private object storage and hand a worker a short-lived signed URL, provided the selected transcription API supports that ingestion pattern. Do not assume it does.

The upload function needs a fresh FormData body for every attempt. Reusing a consumed body is a common retry trap. It also needs a deadline that actually aborts the request rather than merely rejecting a separate timer.

function retryAfterMs(value: string | null): number | undefined {
  if (!value) return undefined;

  const seconds = Number(value);
  if (Number.isFinite(seconds) && seconds >= 0) return seconds * 1_000;

  const dateMs = Date.parse(value);
  if (Number.isNaN(dateMs)) return undefined;
  return Math.max(0, dateMs - Date.now());
}

function delay(ms: number): Promise<void> {
  return new Promise((resolve) => setTimeout(resolve, ms));
}

async function uploadWithBackoff(
  bytes: Uint8Array,
  filePath: string,
  config: UploadConfig,
): Promise<TranscriptResult> {
  for (let attempt = 1; attempt <= config.maxAttempts; attempt += 1) {
    const form = new FormData();
    const name = filePath.split("/").at(-1) ?? "recording.bin";
    form.set("file", new Blob([bytes]), name);

    let response: Response;
    try {
      response = await fetch(config.endpoint, {
        method: "POST",
        headers: { Authorization: `Bearer ${config.apiKey}` },
        body: form,
        signal: AbortSignal.timeout(config.timeoutMs),
      });
    } catch (error) {
      const message = error instanceof Error ? error.message : "Upload failed";
      return { ok: false, stage: "transport", message, retryable: false };
    }

    if (response.ok) {
      const value: unknown = await response.json();
      if (
        typeof value === "object" &&
        value !== null &&
        "text" in value &&
        typeof value.text === "string"
      ) {
        return { ok: true, text: value.text };
      }
      return {
        ok: false,
        stage: "response",
        message: "The provider returned no transcript text.",
        retryable: false,
      };
    }

    const transient = response.status === 429 || response.status >= 500;
    if (!transient || attempt === config.maxAttempts) {
      const detail = await response.text();
      return {
        ok: false,
        stage: "provider",
        message: `Transcription returned ${response.status}: ${detail}`,
        retryable: transient,
      };
    }

    const serverDelay = retryAfterMs(response.headers.get("retry-after"));
    const exponentialDelay = 500 * 2 ** (attempt - 1);
    await delay(serverDelay ?? exponentialDelay);
  }

  return {
    ok: false,
    stage: "provider",
    message: "Transcription attempts were exhausted.",
    retryable: false,
  };
}
Enter fullscreen mode Exit fullscreen mode

This is deliberately conservative: a client timeout does not trigger an automatic retry because the server may still be processing the first request. Without a provider-supported idempotency contract, another multipart POST may duplicate work. A short client deadline plus a visible manual retry is safer than an aggressive loop.

The automatic backoff is limited to 429 and 5xx responses. It honors Retry-After, then falls back to exponential delays of 500, 1,000, and 2,000 milliseconds as attempts increase. Do not retry authentication errors, rejected file types, configured size failures, or an unavailable-capability response. Those require a different input, configuration, or provider.

Choose a provider on correctness, not route resemblance

For this workflow, the winning demo is not a transcript that sounds plausible. It is a transcript that lets the next stage recover invoice fields correctly and route uncertain values to review. Use the same representative recordings, languages, background noise, and supplier-name vocabulary for every candidate.

Option What I would verify first Boundary for this build
OpenAI speech to text Current upload limits, supported formats, and timestamp behavior in its audio guide Evaluate the documented audio workflow before coupling extraction to its response shape
Deepgram Prerecorded-audio request modes and the status/error contract Confirm how the chosen mode handles long files and retries
AssemblyAI Upload and asynchronous transcription lifecycle Budget for job polling and map terminal job states explicitly
Google Cloud Speech-to-Text Long-running recognition and storage requirements Account for cloud-project and object-storage integration work
This runtime Capability readiness from the self-describing discovery surface Speech-to-text is currently unavailable, so it is not eligible for production selection

This is not a universal ranking. It is a shortlist for a one-person SaaS where integration time competes directly with product work. I would run a small acceptance corpus through ready candidates, compare exact field recovery after extraction, and retain the audio-to-transcript boundary so the provider can change without rewriting invoice review. The main limitation of the synchronous options is that a large multipart request keeps more of the transport path coupled to one operation; an asynchronous job API asks for more state handling but gives long work an explicit lifecycle.

The next extraction stage has a different competitor set. Claude from Anthropic, Gemini, OpenRouter, and Together are real options to evaluate alongside OpenAI-compatible model access. They are not substitutes for the speech service merely because they can participate after transcription. Compare them on JSON-schema behavior, the invoice corpus, data-region requirements, and operational fit; then keep that choice behind a separate extraction interface. Mixing these two decisions makes a polished transcript look like proof of correct invoice fields. It is not.

Structured output correctness is the decision axis. Measure whether invoice number, supplier, dates, currency, subtotal, tax, and total survive the whole pipeline. Keep transcript confidence and extraction validation separate; arithmetic checks such as subtotal plus tax matching total can catch a useful class of errors, but they cannot prove that the speaker said the value.

What I would change at scale

The synchronous example is enough to expose the control flow. It is not the final architecture for hour-long recordings.

I would upload once to private storage, enqueue a transcription job with a stable recording ID, and let a worker call the selected provider. The worker would persist state transitions such as uploaded, transcribing, extracting, needs_review, and complete. Consumer-side idempotency would make redelivery harmless. The browser would poll my job resource, not hold one request open while audio crosses two networks and a model runs.

I would also keep raw provider responses behind a narrow adapter. That costs a little code now. It buys back future shipping time when a provider changes or a second provider performs better on noisy supplier calls.

One constraint stays fixed: retries cannot manufacture availability. Check readiness before accepting a recording, enforce your byte gate before upload, and use backoff only after a transient response proves the request reached a service that may recover.

Three gates. One honest fallback.

Further reading

Top comments (0)