A property manager needs a correctly routed maintenance ticket, not a retry loop that eventually returns some text. The least complex workable design is one durable queue, one shared admission gate, and one validator between transcription and triage. Honor Retry-After when a speech-to-text API returns 429, add bounded jitter when that signal is absent, and measure structured-output correctness separately from transport success.
Short answer: treat 429 as a scheduling signal. Do not let every worker sleep and retry independently. Queue each recording once, pause admission at the destination level, and preserve the recording and ticket identifiers through every attempt. Then a burst of elevator-noise reports becomes delayed work rather than duplicated or misrouted work.
| Check | Signal to record | Pick this response when | Failure it prevents |
|---|---|---|---|
| 1. Admission |
retry_after_ms, eligible_at, destination |
The response supplies Retry-After
|
A synchronized retry wave |
| 2. Queue | queue age, attempt, audio hash | Work must survive a process restart | Lost or duplicate recordings |
| 3. Batch boundary | tenant, property, urgency | Many recordings arrive together | One property consuming every slot |
| 4. Contract | schema outcome, reason code | Transcripts feed automatic triage | Valid HTTP responses creating bad tickets |
| 5. Alert | oldest eligible age and correctness ratio | Operators need an actionable page | Alerting on noisy raw 429 counts |
The table is the field guide. The rest shows where each check belongs, then goes deep on a copy-pasteable TypeScript scheduler.
1. How should a speech-to-text API control 429 rate limits?
Use the server's Retry-After value when it is valid. HTTP defines the field as either a delay in seconds or an HTTP date, while status 429 means the client has sent too many requests in a given amount of time. That combination gives the scheduler a concrete next-eligible time. It does not guarantee that the next request will succeed.
This distinction matters. A per-job delay answers "when may this recording try again?" A destination gate answers "when may any worker send another request?" For a shared API quota, the second question is usually the useful one. Ten workers independently receiving the same delay can otherwise wake together.
Pick a destination-level gate when workers share the same upstream limit. Pick a narrower tenant gate only when limits are actually isolated by tenant. The trade-off is deliberate: a shared gate can delay unrelated properties, but it also stops them from amplifying a destination-wide limit. The correct scope comes from the quota contract, not from queue topology.
Parse both legal field forms. Reject invalid or past values, then fall back to exponential delay with jitter and a cap. Keep attempt limits too. Infinite retries can turn an obsolete voicemail into permanent queue pressure.
No header? Use the fallback.
For the example below, that means a one-second base and a 60-second cap. Those values are policy inputs, not universal defaults; choose them against the upstream contract and the maximum delay a maintenance desk can tolerate. Full jitter spreads attempts across the entire calculated window. The explicit trade-off is slower best-case recovery in exchange for fewer workers colliding at the window edge.
2. Why does a durable queue beat sleeping workers?
A sleeping Node.js process owns time in memory. A queued job owns intent in durable state.
That difference is decisive.
For each recording, store an immutable job ID, tenant ID, property ID, audio object reference, audio hash, attempt count, and eligibleAt. Claim eligible jobs with a lease. On success, write the transcript against the same job ID; on 429, release the lease and move eligibleAt; on a terminal result, retain a reason code for review. These are architecture fields, not a vendor-specific queue recipe.
Do not enqueue the audio bytes repeatedly. Keep one durable object and pass its reference. An idempotency key derived from the stable job ID prevents your own completion handler from creating two downstream tickets when a response is delivered twice. Idempotency on the speech request itself depends on the API contract, so do not assume it exists.
Consider three recordings arriving within the same minute: a burst pipe, a broken lobby light, and a resident asking about parking. The first request receives 429 with a 20-second delay. If each worker owns its own timer, the other two can still hit the same constrained destination, receive the same response, and wake beside the first worker. With a shared gate, all three jobs remain durable, the burst-pipe job retains its urgency, and no worker sends until the gate reopens. Fair dequeueing can then select the urgent job first without discarding the other two. This is why retry state, business priority, and transport concurrency must be separate fields.
The useful dashboard is a diagram in words: recording accepted -> durable job -> admission gate -> transcription -> schema validator -> triage queue. Put a counter and a duration at every arrow. Now a red chart can say where flow stopped instead of merely announcing that errors exist.
3. How should a Node.js retry scheduler behave?
Keep policy separate from the HTTP client. The following implementation parses Retry-After, computes bounded full jitter when the header cannot be used, and updates one shared gate. All delays are milliseconds inside the program; that naming choice avoids a surprisingly common seconds-versus-milliseconds mistake.
type Clock = () => number;
type GateState = { blockedUntilMs: number };
const gate: GateState = { blockedUntilMs: 0 };
function parseRetryAfterMs(value: string | null, nowMs: number): number | null {
if (value === null) return null;
const seconds = Number(value);
if (Number.isFinite(seconds) && seconds >= 0) {
return Math.ceil(seconds * 1_000);
}
const dateMs = Date.parse(value);
if (!Number.isFinite(dateMs)) return null;
return Math.max(0, dateMs - nowMs);
}
function fallbackDelayMs(attempt: number, random: () => number): number {
const baseMs = 1_000;
const capMs = 60_000;
const ceilingMs = Math.min(capMs, baseMs * 2 ** attempt);
return Math.floor(random() * ceilingMs);
}
function blockGate(
retryAfter: string | null,
attempt: number,
now: Clock = Date.now,
random: () => number = Math.random,
): number {
const nowMs = now();
const headerDelayMs = parseRetryAfterMs(retryAfter, nowMs);
const delayMs = headerDelayMs ?? fallbackDelayMs(attempt, random);
gate.blockedUntilMs = Math.max(gate.blockedUntilMs, nowMs + delayMs);
return delayMs;
}
async function waitForAdmission(now: Clock = Date.now): Promise<void> {
const delayMs = Math.max(0, gate.blockedUntilMs - now());
if (delayMs === 0) return;
await new Promise<void>((resolve) => setTimeout(resolve, delayMs));
}
The in-memory gate makes the algorithm visible, but it is not the production state store when several processes consume the queue. Put blockedUntilMs in shared storage and update it atomically so an older, shorter delay cannot overwrite a newer, longer one. The Math.max rule is the key invariant.
Wire the policy into one attempt, not an unbounded loop:
type TranscriptionJob = {
id: string;
tenantId: string;
propertyId: string;
audioUrl: string;
attempt: number;
};
type AttemptResult =
| { kind: "completed"; transcript: string }
| { kind: "reschedule"; eligibleAtMs: number; reason: "rate_limited" }
| { kind: "review"; reason: string };
type SendSpeechRequest = (job: TranscriptionJob) => Promise<Response>;
async function transcribeOnce(
job: TranscriptionJob,
sendSpeechRequest: SendSpeechRequest,
): Promise<AttemptResult> {
await waitForAdmission();
const response = await sendSpeechRequest(job);
if (response.status === 429) {
const delayMs = blockGate(response.headers.get("retry-after"), job.attempt);
return {
kind: "reschedule",
eligibleAtMs: Date.now() + delayMs,
reason: "rate_limited",
};
}
if (!response.ok) {
return { kind: "review", reason: `speech_http_${response.status}` };
}
const payload: unknown = await response.json();
if (!hasTranscript(payload)) {
return { kind: "review", reason: "invalid_transcript_response" };
}
return { kind: "completed", transcript: payload.text };
}
function hasTranscript(value: unknown): value is { text: string } {
return (
typeof value === "object" &&
value !== null &&
"text" in value &&
typeof value.text === "string" &&
value.text.trim().length > 0
);
}
The injected sender owns the chosen API's request shape, endpoint, authentication, idempotency behavior, and error contract. A transport wrapper should expose those differences while the queue state machine stays stable.
Test the scheduler with an injected clock and random function. Use cases should include integer seconds, an HTTP date, malformed input, a past date, two overlapping 429 responses, the maximum delay, a restart, and a duplicated completion. A test that only checks "eventually succeeds" misses the timing and duplication bugs that matter.
4. Pick correctness metrics before throughput metrics
A transcription can return 200 and still be unusable for ticket triage. For property management, define a small downstream contract such as category, propertyId, urgency, summary, and evidence. Validate types and allowed values. Route missing or contradictory fields to review rather than guessing.
The transcript remains evidence, not truth. "Water is pouring through the ceiling at 14 King Street" carries urgency and a location, while background speech may contain another address. A validator can prove shape; it cannot prove that the extracted address matches the caller's property. That requires domain checks against known properties and, for high-impact cases, human confirmation.
Track two denominators. Transport success rate uses attempted speech requests. Structured correctness rate uses completed transcriptions that reach the triage contract. Combining them hides a painful state: perfect API availability with steadily degrading routing.
For operations, record job_id, tenant_id, attempt, queue_age_ms, eligible_at, http_status, retry_after_ms, transcript_schema_result, and triage_result. Avoid recording raw audio or full transcripts in routine logs; those can contain resident names, addresses, phone numbers, and access instructions. Keep sensitive content in its governed data path and use identifiers in telemetry.
Alert on consequences. A rising oldest-eligible queue age means callers are waiting. A falling structured correctness ratio means automation is producing unsafe work. A raw 429 counter is diagnostic context, not necessarily a page: the queue may be absorbing it exactly as designed.
Page on harm, not noise.
5. Limits of this pattern
This design cannot infer an undocumented quota scope. It cannot make a non-idempotent upstream operation idempotent, and it cannot guarantee fair service unless the dequeue policy includes tenant or property fairness. Long recordings may also consume a different resource budget than short ones, so job count alone can be a weak capacity measure.
Start with one shared gate because it is easy to reason about. Split gates only after quota documentation or telemetry shows independent limits. Add weighted admission only when audio duration demonstrably predicts resource use. Each extra scheduler dimension creates state that must be tested during restarts and concurrent updates.
The boundary is crisp: 429 handling protects admission, the durable queue protects work, and schema plus domain validation protects ticket correctness. You need all three before automatic triage deserves trust.
References
- RFC 6585, Additional HTTP Status Codes: https://www.rfc-editor.org/rfc/rfc6585
- RFC 9110, HTTP Semantics (
Retry-After): https://www.rfc-editor.org/rfc/rfc9110 - MDN,
Retry-After: https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Retry-After - Node.js documentation, timers: https://nodejs.org/api/timers.html
Further reading
- MDN, HTTP response status codes: https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status
- OpenTelemetry semantic conventions: https://opentelemetry.io/docs/specs/semconv/
Top comments (0)