Sending a five-minute recording to a speech-to-text API is easy. Sending a two-hour meeting and getting back a transcript that is complete, ordered, and safe to retry is a different problem.
The model is only one part of the system. A production pipeline also has to deal with upload limits, long processing times, chunk boundaries, repeated text, timestamp drift, speaker labels, partial failures, and users refreshing the page while the job is still running.
This article describes a provider-neutral architecture for handling those problems. The examples use TypeScript-like pseudocode, but the same design works with most queues, object stores, databases, and transcription providers.
Start with a job, not an HTTP request
A long transcription should not depend on one browser connection remaining open.
The upload endpoint should store the media, create a job, and return a stable identifier. A worker can then continue even if the user closes the tab.
type TranscriptionJob = {
id: string;
sourceFileKey: string;
status: "queued" | "processing" | "completed" | "failed";
attempt: number;
createdAt: string;
updatedAt: string;
};
async function createJob(file: UploadedFile) {
const sourceFileKey = await objectStore.put(file);
const job = await db.transcriptionJobs.insert({
sourceFileKey,
status: "queued",
attempt: 0,
});
await queue.publish("transcription.requested", { jobId: job.id });
return { jobId: job.id, status: job.status };
}
The UI can poll GET /transcriptions/:jobId or subscribe to status events. Either way, the job state belongs in the backend rather than in the browser.
Normalize the media before transcription
Users upload unpredictable files: stereo video, quiet voice notes, variable frame-rate recordings, and formats that a transcription provider may not accept directly.
A normalization step gives the rest of the pipeline one stable input format.
For spoken audio, a common intermediate representation is mono PCM audio at a fixed sample rate. The exact settings depend on the model and provider, so check their documentation rather than converting every file to an arbitrary format.
ffmpeg -i input.mov \
-vn \
-ac 1 \
-ar 16000 \
-c:a pcm_s16le \
normalized.wav
Normalization is also a useful place to inspect duration, detect a file with no audio stream, and reject corrupt media before spending money on inference.
Store the normalized file separately from the original. The original may be needed for exports or debugging, while the normalized version should be treated as a reproducible processing artifact.
Split on silence, then add overlap
Long recordings usually have to be divided into smaller sections. A fixed split every 30 seconds is simple, but it can cut through a word or sentence.
A better strategy is:
- choose a target chunk length;
- search near that boundary for a suitable silence;
- add a small overlap on both sides;
- keep the exact source start and end time for every chunk.
type AudioChunk = {
index: number;
startMs: number;
endMs: number;
overlapBeforeMs: number;
overlapAfterMs: number;
objectKey: string;
};
The overlap gives the model context at the edges. It also creates duplicate text, which must be removed during the merge step.
This is a known long-form ASR problem. The original Whisper paper describes buffered transcription over consecutive windows, while the OpenAI speech-to-text guide discusses carrying context from a previous chunk and supplying uncommon names or acronyms.
Make every chunk independently retryable
A two-hour job should not restart from zero because chunk 37 failed.
Persist the state of each chunk:
type ChunkResult = {
jobId: string;
chunkIndex: number;
status: "pending" | "processing" | "completed" | "failed";
providerRequestId?: string;
transcriptKey?: string;
errorCode?: string;
attempt: number;
};
Before processing a chunk, the worker checks whether a completed result already exists. If it does, the worker exits without calling the provider again.
async function processChunk(jobId: string, chunk: AudioChunk) {
const existing = await db.chunkResults.find(jobId, chunk.index);
if (existing?.status === "completed") return;
await db.chunkResults.upsert({
jobId,
chunkIndex: chunk.index,
status: "processing",
attempt: (existing?.attempt ?? 0) + 1,
});
try {
const result = await transcriber.transcribe(chunk.objectKey);
const transcriptKey = await objectStore.putJson(result);
await db.chunkResults.complete(jobId, chunk.index, transcriptKey);
} catch (error) {
await db.chunkResults.fail(jobId, chunk.index, classify(error));
throw error;
}
}
Retry temporary failures with exponential backoff and jitter. Do not retry permanent errors such as an unsupported codec forever. A dead-letter queue is useful for jobs that need investigation.
Idempotency matters here. Queue systems commonly deliver a message more than once, and workers can crash after the provider completes but before the database update succeeds. Reprocessing must not silently create duplicate transcript sections or duplicate charges.
Keep timestamps in the source timeline
Most providers return timestamps relative to the beginning of the uploaded chunk. Users need timestamps relative to the complete recording.
The conversion is straightforward:
function toSourceTime(chunk: AudioChunk, localMs: number) {
return chunk.startMs + localMs;
}
The difficult part is the overlap. Two neighboring chunks can both contain the same sentence with slightly different text and timestamps.
Do not simply append every result. A basic merge algorithm should:
- compare words or segments inside the overlap interval;
- estimate the best matching boundary;
- retain one copy of the shared speech;
- reject segments outside the chunk’s trusted region;
- keep timestamps monotonic.
For word-level timestamps, sequence alignment works better than exact string equality because recognition around the boundary may differ by punctuation or one word.
function mergeWithOverlap(
previous: TranscriptSegment[],
next: TranscriptSegment[],
overlap: TimeRange,
) {
const left = segmentsInside(previous, overlap);
const right = segmentsInside(next, overlap);
const boundary = findBestAlignment(left, right);
return [
...segmentsBefore(previous, boundary.leftTime),
...segmentsAfter(next, boundary.rightTime),
];
}
The exact alignment method can evolve. The important architectural choice is to preserve raw chunk results so the transcript can be re-merged later without paying to transcribe the audio again.
Treat speaker labels as provisional
Speaker diarization across independent chunks has a predictable failure mode: Speaker 1 in one chunk may not be the same person as Speaker 1 in the next.
If the provider assigns local speaker labels, store them as local identifiers:
type LocalSpeakerId = `${number}:${string}`;
// Examples: "12:speaker_0", "13:speaker_0"
A later reconciliation step can map local labels to stable job-level speakers using voice embeddings, temporal continuity, participant metadata, or human corrections.
Do not overwrite the original diarization output. Store the mapping separately so corrections remain reversible.
Separate transcription from document generation
The raw transcript and the document shown to the user should not be the same database field.
Keep several layers:
- immutable provider responses;
- merged word or segment timeline;
- speaker mapping;
- formatted transcript;
- summaries, action items, or exports.
This separation makes it possible to fix formatting or speaker labels without retranscribing the file. It also prevents a summary-generation failure from turning a successful transcription job into a failed one.
A simple state machine might look like this:
uploaded
-> normalizing
-> chunking
-> transcribing
-> merging
-> formatting
-> completed
Optional tasks such as summaries and subtitle exports can run after completed and expose their own status.
Measure the complete user experience
Provider latency is only one number. A user experiences the entire path from upload to usable text.
Useful metrics include:
- upload time;
- queue wait time;
- normalization time;
- transcription time per audio minute;
- retry rate by error class;
- percentage of jobs requiring a manual retry;
- merge and formatting time;
- total time to a usable transcript;
- cost per processed audio hour.
Track job size and media type with these metrics. Averages can hide the fact that short MP3 files are fast while long video files consistently fail during normalization.
Quality also needs operational signals. Repeated phrases, empty chunks, non-monotonic timestamps, sudden language changes, and implausibly high words-per-second can all trigger warnings before the user opens a broken result.
Decide what “complete” means
A transcript can be technically complete while still missing the features the user expects.
Define the contract explicitly. For example:
type CompletedTranscript = {
text: string;
segments: TranscriptSegment[];
durationMs: number;
detectedLanguage?: string;
speakers?: Speaker[];
warnings: QualityWarning[];
};
If speaker labels are optional, say so. If the system could not confidently detect the language, return that uncertainty rather than inventing certainty. If one chunk failed, decide whether the product should expose a partial transcript or block the entire result.
These choices shape the product more than the name of the ASR model.
A pipeline is the feature
When we work on products such as TranscribeThis, it is tempting to discuss transcription as one API call. Users never experience that API call in isolation. They experience whether their long recording finishes, whether refreshing the page loses progress, whether timestamps make sense, and whether a failed section can be recovered.
The reliable design is not complicated because long-form speech requires fashionable infrastructure. It is complicated because a long recording creates many opportunities for partial failure.
Build around stable jobs, reproducible media, retryable chunks, preserved raw results, and a deterministic merge. Then treat formatting, summaries, and exports as separate layers.
That architecture gives you something more valuable than a successful demo: a transcription workflow that can survive real files from real users.
- Disclosure: I work on TranscribeThis. This article was prepared with AI assistance and then reviewed and edited by the author, who takes responsibility for the technical content.*
Top comments (0)