Most "auto captions" tools start the same way: upload your video, wait, download the result. I wanted the opposite for short vertical videos (TikTok, Reels, Shorts): the file never leaves the browser. Transcription, editing, rendering and export all run on the user's machine.
The result is a small static site: sous-titres.pages.dev (the UI is in French, but it transcribes 8 languages). This post walks through the architecture, the real code, and the limits I hit.
The pipeline
- Decode the audio track in the page (
AudioContextat 16 kHz, mixed down to mono). - Send the samples to a Web Worker running Whisper through transformers.js, on WebGPU if available, WASM otherwise.
- Turn word-level timestamps into short caption segments (pure functions, unit-tested in Node).
- Draw every frame on a
<canvas>: source video + captions where the current word lights up. - Record the canvas plus the original audio with MediaRecorder and hand the user a file.
No backend touches the media. The only server code is a tiny Cloudflare Pages Function for anonymous counters and license checks.
1. Whisper in a worker, WebGPU first
The worker picks a device and a precision. WebGPU gets fp16 for the encoder when the adapter supports shader-f16, and a 4-bit decoder; the WASM fallback uses 8-bit weights for both:
async function detecterWebGPU() {
try {
if (!self.navigator?.gpu) return { ok: false };
const adaptateur = await navigator.gpu.requestAdapter();
if (!adaptateur) return { ok: false };
return { ok: true, fp16: adaptateur.features.has("shader-f16") };
} catch {
return { ok: false };
}
}
const device = gpu.ok ? "webgpu" : "wasm";
const dtype = gpu.ok
? { encoder_model: gpu.fp16 ? "fp16" : "fp32", decoder_model_merged: "q4" }
: { encoder_model: "q8", decoder_model_merged: "q8" };
Two models are offered: onnx-community/whisper-base_timestamped (~90 MB) and whisper-small_timestamped (~250 MB). The _timestamped exports matter: they expose the cross-attention needed for return_timestamps: "word".
navigator.gpu existing is not enough: an adapter can be returned and the pipeline can still fail to build on it. So creation is wrapped and retried in WASM:
try {
({ transcripteur: t, device } = await obtenirTranscripteur(modele, false));
} catch (e) {
self.postMessage({ type: "info", message: `WebGPU indisponible (${e.message || e}), passage en mode compatible.` });
({ transcripteur: t, device } = await obtenirTranscripteur(modele, true));
}
There is a second fallback: if word-level timestamps throw, the worker asks for sentence-level timestamps and the UI spreads the words across each sentence, weighted by length.
2. From Whisper chunks to captions
Whisper's word output is messier than it looks. Tokens sometimes arrive split (" micro" + "-entrepreneur."), and noise markers like [Musique] show up. The rule I settled on: a chunk that does not start with a space belongs to the previous word.
const colle = precedent && !/^\s/.test(brut) && !/^[\[(]/.test(texte);
if (colle) {
precedent.texte += texte;
precedent.fin = Math.max(precedent.fin, fin);
} else {
mots.push({ texte, debut, fin });
}
Segmentation is a greedy pass with a few cut conditions: max words per line (user setting, 1 to 5), max 22 characters, a pause longer than 0.6 s, end of sentence, or a comma once the line has two words. Each segment stays on screen 0.3 s after its last word, without ever overlapping the next one. All of this lives in a DOM-free module, so it runs under node --test and also generates the .srt and the .ass karaoke (\k tags) exports.
When the user fixes a typo, the segment keeps its start and end times and the new words are spread inside them. Small, but it means editing never breaks the sync.
3. Rendering: one function, three styles
The same dessinerSousTitres(ctx, segments, t, settings) draws the live preview and the export. It measures each word, wraps at 86% of the 1080 px width, then draws word by word with a thick black stroke, a soft shadow, and a short "pop" on the active word:
const age = t - seg.mots[i].debut;
const pop = estActif ? 1 + 0.1 * Math.max(0, 1 - age / 0.18) : 1;
ctx.save();
ctx.translate(x + w / 2, y);
ctx.scale(pop, pop);
Three styles are just branches in that loop: highlight the active word's color, draw a rounded box behind it, or reveal words only once they've been spoken. The font (Montserrat 900) is self-hosted and preloaded, and the app awaits document.fonts.load() before the first draw. Otherwise the first frames get measured with a fallback font.
Because this module only needs a 2D context, I reused it outside the browser: the promo clip for the tool was rendered in Node with @napi-rs/canvas, the same rendu.js, and transformers.js on CPU.
4. Export: canvas + MediaRecorder
The export plays the video in a hidden <video>, redraws each frame on requestAnimationFrame, and records canvas.captureStream(30) plus the original audio routed through createMediaElementSource:
const candidats = [
"video/mp4;codecs=avc1.640028,mp4a.40.2",
"video/mp4;codecs=avc1,mp4a.40.2",
"video/mp4",
"video/webm;codecs=vp9,opus",
"video/webm;codecs=vp8,opus",
"video/webm",
];
return candidats.find((m) => MediaRecorder.isTypeSupported(m)) || null;
Recent Chrome, Edge and Safari produce MP4; Firefox produces WebM. The honest trade-off: export runs in real time (60 s of video takes about 60 s) and the tab must stay visible, because background tabs throttle rAF. WebCodecs plus a JS muxer would allow faster-than-real-time export and more consistent output. That's the next step, and it isn't shipped yet.
5. The deployment surprise: Hugging Face and *.workers.dev
I first served the app from a *.workers.dev subdomain, and the model never loaded there: model file requests answered 404 whenever the Referer was a workers.dev origin. Same files, same code, different referrer.
Rather than fight it, I moved the static app to Cloudflare Pages (*.pages.dev), where the model loads normally, and left a one-line Worker on the old address that answers 301 to the new one:
const CIBLE = "https://sous-titres.pages.dev";
export default {
fetch(requete) {
const url = new URL(requete.url);
return Response.redirect(CIBLE + url.pathname + url.search, 301);
},
};
Lesson: if your app fetches third-party assets client-side, test from the real production origin early. The Referer your users send is part of your dependency surface.
Privacy, verifiable
"No upload" is easy to claim, so the FAQ tells users how to check it: open DevTools → Network, run a transcription, and you'll see only the model (Hugging Face), the library (jsDelivr) and anonymous event counters (a visit, an export). No cookies, no third-party analytics.
Limits (the real ones)
- Desktop Chrome/Edge recommended (that's where I tested it). Firefox and Safari aren't tested yet; without WebGPU the app falls back to WASM, which is slower. Phones can run out of memory.
- First run downloads the model (~90 MB or ~250 MB), then it's cached.
- Speed depends on hardware. With WebGPU, a one-minute clip usually transcribes in under a minute. Without it, expect several minutes.
- Whisper makes mistakes on names, numbers and acronyms. The editor exists for that.
- Short-form only. Beyond 3–5 minutes, browser memory becomes the bottleneck.
- Real-time export, as described above.
Try it
sous-titres.pages.dev: free, no account, with a small watermark on exported videos. A one-time paid license removes it. Feedback on the WebGPU path, especially from AMD and Intel GPUs, is very welcome in the comments.
Top comments (0)