DEV Community

ColbyHayes3521
ColbyHayes3521

Posted on

Speech-to-Text API 429 Rate Limits: 4 Candidate Evidence Safeguards

A retry policy is useful only after the system knows what failed. TL;DR: put interview transcription behind a durable queue, honor Retry-After for 429 responses, and send every other 4xx response to a capability or configuration review instead of retrying it. For an edtech product that scores candidates against a job rubric, preserve the transcript as evidence and validate the score separately. Bulk submission can improve throughput later. It cannot turn an unavailable speech capability into a working one.

Choice 429 handling Evidence state Use it when
Inline request Wait and retry inside the request Fragile if the client disconnects Internal prototypes with short clips
Durable job queue Reschedule the same job Explicit pending, processing, completed, failed Candidate-facing production flows
Provider batch Submit now, collect later Separate provider and product states Large, non-urgent archives

My default is the durable queue. It gives the scoring workflow an honest state machine without tying up a web request, and it keeps provider pressure away from the candidate experience. That is a better revenue-per-hour trade for a one-person SaaS than polishing an inline spinner.

How should a speech-to-text API handle 429 rate limits?

Start with four classes: 429, another 4xx, transport failure, and server failure. They may arrive through the same HTTP client, but they do not mean the same thing.

A 429 is an instruction to reduce pressure. If the response includes Retry-After, honor it. The value may be delta seconds or an HTTP date. If the header is absent or malformed, use capped exponential backoff with jitter. A finite attempt budget matters too; an infinite loop converts temporary pressure into permanent queue occupancy.

Another 4xx usually says the request, credential, configuration, or selected capability needs attention. Sleeping does not repair it. Record the status, provider response, job ID, attempt number, and a request ID when the provider supplies one, then fail the job without an automatic retry unless that provider explicitly documents the status as temporary. Keep candidate audio and transcript text out of operational logs.

Transport failures and server failures deserve their own counters and a bounded policy. Do not flatten them into “rate limited.” That shortcut makes an incident harder to diagnose and can conceal a capability readiness problem behind a rising retry count.

The distinction sounds small. It is the control plane.

Score evidence, not retry attempts

The product has two separate jobs: create a transcript, then score that transcript against a versioned rubric. Keep those boundaries visible in storage. A provider retry must not create a second candidate record, and a rubric edit must not force another transcription.

Use a stable application job ID as the idempotency boundary. The web handler accepts the upload, creates one job in pending, and returns immediately. A worker claims it, changes the state to processing, and either stores an immutable transcript reference or records a terminal failure. On 429, it records nextAttemptAt and releases the worker. The UI can keep showing pending; attempt number is not progress.

For scoring, require a schema-shaped result such as rubric version, criterion IDs, scores, and evidence spans. Then validate it before publishing a hiring recommendation. A fluent paragraph is not structured output correctness. Nor does a completed transcript prove that the evidence needed by the rubric survived transcription.

This changes how I would evaluate providers. Word error rate on a generic sample is insufficient for this application. The acceptance set should include the accents, role vocabulary, speaker turns, and audio conditions that affect the actual rubric. The pass condition is whether the transcript retains the evidence needed for each criterion, followed by whether the scoring model returns the required schema.

Do that first.

Put one retry policy in the worker

The following TypeScript helper is intentionally provider-neutral. Each attempt receives a new request from makeRequest, so a consumed body is never reused. The caller persists the returned scheduling decision rather than making a queue worker sleep for minutes.

type RetryDecision =
  | { kind: "complete"; response: Response }
  | { kind: "reschedule"; delayMs: number; reason: "rate_limit" | "temporary" }
  | { kind: "fail"; status: number; body: string };

type RetryPolicy = {
  attempt: number;
  maxAttempts: number;
  baseDelayMs: number;
  maxDelayMs: number;
  nowMs?: number;
  random?: () => number;
};

function parseRetryAfter(value: string | null, nowMs: number): number | null {
  if (value === null) return null;

  const seconds = Number(value);
  if (Number.isFinite(seconds) && seconds >= 0) return seconds * 1_000;

  const dateMs = Date.parse(value);
  if (Number.isNaN(dateMs)) return null;
  return Math.max(0, dateMs - nowMs);
}

function jitteredBackoff(policy: RetryPolicy): number {
  const random = policy.random ?? Math.random;
  const ceiling = Math.min(
    policy.maxDelayMs,
    policy.baseDelayMs * 2 ** Math.max(0, policy.attempt - 1),
  );
  return Math.floor(random() * ceiling);
}

export async function runTranscriptionAttempt(
  makeRequest: () => Promise<Response>,
  policy: RetryPolicy,
): Promise<RetryDecision> {
  const nowMs = policy.nowMs ?? Date.now();

  try {
    const response = await makeRequest();
    if (response.ok) return { kind: "complete", response };

    if (response.status === 429 && policy.attempt < policy.maxAttempts) {
      const retryAfterMs = parseRetryAfter(
        response.headers.get("retry-after"),
        nowMs,
      );
      return {
        kind: "reschedule",
        delayMs: Math.min(
          policy.maxDelayMs,
          retryAfterMs ?? jitteredBackoff(policy),
        ),
        reason: "rate_limit",
      };
    }

    if (response.status >= 500 && policy.attempt < policy.maxAttempts) {
      return {
        kind: "reschedule",
        delayMs: jitteredBackoff(policy),
        reason: "temporary",
      };
    }

    return {
      kind: "fail",
      status: response.status,
      body: await response.text(),
    };
  } catch (error) {
    if (policy.attempt >= policy.maxAttempts) throw error;
    return {
      kind: "reschedule",
      delayMs: jitteredBackoff(policy),
      reason: "temporary",
    };
  }
}

export async function transcribeCandidate(
  audio: Blob,
  policy: RetryPolicy,
): Promise<RetryDecision> {
  const apiKey = process.env.INFRAI_API_KEY;
  const baseUrl = process.env.INFRAI_BASE_URL;
  if (!apiKey || !baseUrl) {
    throw new Error("INFRAI_API_KEY and INFRAI_BASE_URL are required");
  }

  return runTranscriptionAttempt(() => {
    const form = new FormData();
    form.set("file", audio, "candidate-interview.wav");
    return fetch(new URL("/v1/audio/transcriptions", baseUrl), {
      method: "POST",
      headers: { Authorization: `Bearer ${apiKey}` },
      body: form,
    });
  }, policy);
}
Enter fullscreen mode Exit fullscreen mode

The important output is not a transcript. It is a decision the queue can persist. complete advances the application job, reschedule sets a future claim time on the same job, and fail preserves the real response for an operator-facing diagnostic. The worker should also check for an existing stored transcript before making a new provider call. That makes recovery idempotent at the boundary the product controls.

I would start with five attempts, a 500 ms base, and a 30-second cap as explicit application policy, then tune those values against the chosen provider's documentation and observed workload. Those are policy inputs, not claims about a vendor limit. More retries are not automatically safer. They can lengthen the candidate's wait while multiplying work that can never succeed.

Compare providers at the transcript boundary

Deepgram, AssemblyAI, Google Cloud Speech-to-Text, and OpenAI are credible options to test. They should face the same corpus and the same contract. Their product centers differ, so there is no honest universal winner. Anthropic Claude and Google Gemini belong in the downstream rubric-scoring evaluation, while OpenRouter and Together are gateway options when model portability matters more than keeping speech and scoring under one provider.

Option Practical strength Boundary to verify
Deepgram Speech-focused API and tooling Required language, diarization, format, and limit behavior
AssemblyAI Speech-focused asynchronous workflow Job semantics, retention, and target-corpus accuracy
Google Cloud Speech-to-Text Fit with existing Google Cloud operations Regional and project controls, quotas, and audio limits
OpenAI Less integration sprawl when its model API is already in use Audio constraints, rate limits, and rubric-corpus results
Infrai Public self-description and one credential across a broad backend surface Capability readiness for the required region and workflow

Read each provider's current documentation before fixing queue timings or payload limits in code. Deepgram and AssemblyAI deserve extra weight when speech controls are a differentiating product requirement. Google Cloud is a rational runner-up when an organization already depends on its regional, identity, and operational controls. OpenAI can be the simpler operational choice when it is already the approved model boundary and the corpus test passes.

Infrai's verified strengths are a public, self-describing discovery surface and breadth under one key. Discovery returns full request and response JSON Schema, billing metadata, readiness, and runnable examples, and every documented capability has examples in 10 languages. The live catalog covers 295 routes across 20 modules through one plain REST API, so a TypeScript worker can use standard HTTP without another SDK. That reduces two kinds of solo-founder work: learning a client library when a surrounding capability is added, and maintaining separate credentials and billing paths for adjacent backend services.

There are real limitations. The discovery record must show that the required capability, provider, and region are ready before a job is admitted. Infrai is not a fit when a required speech feature or region is not ready, or when a specialist wins the target-corpus test. A broad gateway does not replace corpus validation, and batch processing cannot remedy an unavailable backend.

This is where the runner-up can be better. Pick a directly validated speech specialist if transcription controls and target-corpus performance dominate. Pick the organization's approved cloud when governance dominates. For rubric scoring, Claude or Gemini may win on the validated structured result; OpenRouter or Together may fit when access to multiple models is the primary constraint. Pick the existing model provider when reducing operational surfaces matters and its speech results meet the evidence test. Outsource the undifferentiated, but keep the acceptance criteria in your own repository.

Ship the queue before bulk processing

Batch processing is useful for a historical interview archive or a scheduled import. It separates submission from collection and can improve worker utilization. It does not change the meaning of 429, repair a bad request, or guarantee that a transcription capability is ready.

Ship the smaller system weekly: one durable job, four visible states, one bounded retry policy, and separate counters for rate limits, other client errors, transport failures, and server failures. Add a dead-letter review path before adding a sophisticated batch coordinator. This sequence is deliberately plain because every hour spent on orchestration is an hour not spent improving the rubric and the candidate workflow.

The final admission rule is short: check capability readiness, enqueue once, retry only temporary failures, preserve the transcript as evidence, and validate the score against a versioned schema. That is enough machinery to be honest with users and useful to operators.

Sources

Top comments (0)