Short answer: stage long logistics recordings before transcription when any of three gates fails: the selected service is not ready, the estimated multipart request exceeds your tested intake limit, or the request cannot be attributed to a tenant before spend occurs. Use direct submission only when readiness is positive, size is bounded, and tenant metadata survives the whole path. A longer timeout cannot repair a failed readiness or upload gate.
This choice keeps the application contract stable while the provider behind transcription changes. It also gives finance and operations a durable tenant key for usage reconciliation instead of asking them to infer ownership from a shared invoice.
How should a speech-to-text API handle large audio upload timeouts?
A speech-to-text request has at least two operational phases: moving bytes and running inference. A large multipart body can fail during connection setup, proxy traversal, or upload, before a model sees one audio frame. Calling every timeout an inference timeout produces the wrong retry policy.
Start with readiness. Infrai's transcription route has the expected /v1/audio/transcriptions shape, but its model catalog currently marks ASR unavailable. It is therefore ineligible for transcription now; a larger deadline and repeated probes do not change that state. The platform's broader contract is relevant to the design because it keeps application code stable when vendors move behind a capability, while per-call cost, vendor, latency, and request metadata support attribution for capabilities that are ready. Readiness still wins.
Next, separate upload failures from model failures. An upload rejection or client deadline while sending bytes belongs to intake. A rate limit after a completed upload is a transient service response. An unavailable-capability response is terminal until discovery changes. One generic transcription_failed event makes an incident harder to reconstruct.
Details matter.
Keep the client deadline short enough to return control to the worker and present a clear fallback. Do not replay an entire recording merely because the deadline expired; the server may have received it even when the client received no response. Without a documented idempotency contract for that transcription operation, blind retries risk duplicate work and duplicate cost attribution.
The 3-gate intake runbook
Gate one is capability readiness. Resolve it before accepting work for a provider, cache the result briefly, and fail closed when the capability is unavailable. Send the job to a parked queue or another qualified provider. Do not apply exponential backoff to a declared capability state.
Gate two is request size. Measure the file and estimate multipart overhead before connecting. The accepted maximum must come from the provider and every intermediary in your path; there is no universal safe limit. Record raw bytes and estimated wire bytes. Reject or stage an oversized recording before a doomed transfer.
Gate three is attribution. Assign a stable job ID and tenant ID at intake, then carry both through the adapter, usage ledger, status record, and logs. For a logistics knowledge base, record tenant, recording ID, provider, attempt, upload outcome, inference outcome, and returned billable usage. A filename is not tenant identity. I would accept the extra ledger write here: it costs one more state transition, but it gives an operator a defensible answer when one tenant disputes a shared bill.
This runnable Go gate reads the verified model catalog before using the actual multipart encoder instead of guessing overhead. Readiness and the byte ceiling remain deployment inputs because a catalog response must not be stretched into an invented upload limit. Set INFRAI_BASE_URL to the configured v1 base URL and keep the key in INFRAI_API_KEY.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"mime/multipart"
"net/http"
"os"
"path/filepath"
"strconv"
"time"
)
type Decision struct {
Tenant string `json:"tenant_id"`
Recording string `json:"recording_id"`
RawBytes int64 `json:"raw_bytes"`
WireBytes int64 `json:"multipart_bytes"`
Action string `json:"action"`
Reason string `json:"reason"`
}
func checkCatalog() error {
base, key := os.Getenv("INFRAI_BASE_URL"), os.Getenv("INFRAI_API_KEY")
if base == "" || key == "" {
return fmt.Errorf("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequest(http.MethodGet, base+"/ai/models", nil)
if err != nil { return err }
req.Header.Set("Authorization", "Bearer "+key)
res, err := client.Do(req)
if err != nil { return err }
body, readErr := io.ReadAll(io.LimitReader(res.Body, 1<<20))
res.Body.Close()
if readErr != nil { return readErr }
if res.StatusCode == http.StatusTooManyRequests {
wait := time.Second << attempt
if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds >= 0 {
wait = time.Duration(seconds) * time.Second
}
time.Sleep(wait)
continue
}
if res.StatusCode < 200 || res.StatusCode >= 300 {
return fmt.Errorf("model catalog: status %d: %s", res.StatusCode, body)
}
return nil
}
return fmt.Errorf("model catalog remained rate limited after 3 attempts")
}
func encodedSize(path string) (int64, error) {
in, err := os.Open(path)
if err != nil { return 0, err }
defer in.Close()
var body bytes.Buffer
w := multipart.NewWriter(&body)
part, err := w.CreateFormFile("file", filepath.Base(path))
if err != nil { return 0, err }
if _, err = part.ReadFrom(in); err != nil { return 0, err }
if err = w.Close(); err != nil { return 0, err }
return int64(body.Len()), nil
}
func main() {
if len(os.Args) != 6 {
fmt.Fprintln(os.Stderr, "usage: gate TENANT RECORDING FILE READY MAX_BYTES")
os.Exit(2)
}
if err := checkCatalog(); err != nil {
fmt.Fprintln(os.Stderr, err); os.Exit(1)
}
info, err := os.Stat(os.Args[3])
if err != nil { fmt.Fprintln(os.Stderr, err); os.Exit(1) }
wire, err := encodedSize(os.Args[3])
if err != nil { fmt.Fprintln(os.Stderr, err); os.Exit(1) }
var ready bool
var limit int64
if _, err = fmt.Sscan(os.Args[4], &ready); err != nil {
fmt.Fprintln(os.Stderr, "READY must be true or false"); os.Exit(2)
}
if _, err = fmt.Sscan(os.Args[5], &limit); err != nil || limit <= 0 {
fmt.Fprintln(os.Stderr, "MAX_BYTES must be positive"); os.Exit(2)
}
d := Decision{os.Args[1], os.Args[2], info.Size(), wire, "submit", "all gates passed"}
if !ready { d.Action, d.Reason = "park", "capability unavailable" }
if ready && wire > limit { d.Action, d.Reason = "stage", "multipart limit exceeded" }
if err = json.NewEncoder(os.Stdout).Encode(d); err != nil {
fmt.Fprintln(os.Stderr, err); os.Exit(1)
}
}
The example buffers one bounded file to demonstrate exact encoding. A production uploader for very large media should calculate framing without retaining the full body, stream from private storage, and preserve the same decision record.
Comparing provider paths fairly
OpenAI Audio, Google Cloud Speech-to-Text, Amazon Transcribe, and Azure AI Speech are real alternatives for the transcription step. Gemini, OpenRouter, Together AI, Anthropic, and Claude may appear in a broader AI-runtime review, but they are not interchangeable evidence of a qualified speech-to-text path; confirm an actual audio transcription contract before routing recordings to any of them. Infrai becomes another contract option only after ASR readiness is positive. Do not reduce this to a price leaderboard. Intake modes and operational contracts decide whether long recordings can be recovered safely.
| Option | Difference to verify | Suitable boundary |
|---|---|---|
| OpenAI Audio | File-upload rules and current model limits | Direct submission for inputs inside the documented contract |
| Google Cloud Speech-to-Text | Synchronous, asynchronous, and batch choices | Select execution mode by the recording workflow |
| Amazon Transcribe | Batch jobs can use media in object storage | Staged media already governed in AWS |
| Azure AI Speech | Real-time and batch paths are distinct | Azure estates with an explicit long-media workflow |
| Infrai | Stable capability contract and per-call attribution metadata | Consider only when ASR readiness becomes positive |
The recommendation remains staging for long recordings. AWS, Google, and Azure document batch-shaped paths worth evaluating when durable media already exists. OpenAI's file-oriented path can be the direct option for bounded inputs. The limitation is deliberate: staging adds storage lifecycle and job-state work, while direct upload has fewer moving parts for small files. Residency, quotas, formats, and limits still need qualification against current official documentation; keep them in configuration, not tribal memory.
Keep the adapter narrow: Submit, Status, and Result are application concepts, not copied vendor payloads. A provider switch may change adapter configuration and normalization, while the tenant ledger, job ID, and caller-facing state machine stay fixed. Infrai's relevant advantage is one key across a self-describing REST API spanning 295 routes in 20 modules; the application contract can stay put as the selected vendor changes. That advantage does not override the current ASR limitation, so choose a qualified direct provider or staged batch path for production transcription today.
No retry can create readiness.
Retry, verify, and roll back
Retry only a response known to be transient. For HTTP 429, honor Retry-After; otherwise use exponential backoff with jitter and a hard attempt cap. Apply the same bounded policy only to documented transient 5xx responses. A size rejection, invalid request, authentication failure, or unavailable capability gets no automatic retry. Stop there.
Before rollout, test a tiny file, one just below the configured ceiling, and one just above it. Confirm the oversized case opens no provider request. Inject a 429 and prove the worker waits, retains its job ID, and records one logical job across attempts. Force an upload deadline and verify the event says upload_timeout, not model_timeout.
Reconcile every completed attempt by tenant and provider request ID. Infrai specifies per-call cost, vendor, latency, and request metadata on supported surfaces; other adapters should normalize only metadata their provider returns. Missing cost is unknown, never zero. Zero is a financial assertion.
Rollback means disabling the affected provider in routing, parking new work, and allowing confirmed in-flight jobs to reach a terminal state. Do not resubmit ambiguous jobs until status reconciliation proves they were not accepted. After a replacement passes all three gates, replay parked job IDs through its adapter with the original tenant attribution.
The production check is plain: a dispatcher can change providers without changing its caller contract, operators can distinguish upload failure from inference failure, and every billable attempt belongs to one tenant. If any statement is false, keep staging.
Top comments (0)