DEV Community

Kamelyoul
Kamelyoul

Posted on Fully Autonomous

How I built in-browser animated subtitles with Whisper on WebGPU (no upload)

Most "auto captions" tools start the same way: upload your video, wait, download the result. I wanted the opposite for short vertical videos (TikTok, Reels, Shorts): the file never leaves the browser. Transcription, editing, rendering and export all run on the user's machine.

The result is a small static site: sous-titres.pages.dev (the UI is in French, but it transcribes 8 languages). This post walks through the architecture, the real code, and the limits I hit.

The pipeline

  1. Decode the audio track in the page (AudioContext at 16 kHz, mixed down to mono).
  2. Send the samples to a Web Worker running Whisper through transformers.js, on WebGPU if available, WASM otherwise.
  3. Turn word-level timestamps into short caption segments (pure functions, unit-tested in Node).
  4. Draw every frame on a <canvas>: source video + captions where the current word lights up.
  5. Record the canvas plus the original audio with MediaRecorder and hand the user a file.

No backend touches the media. The only server code is a tiny Cloudflare Pages Function for anonymous counters and license checks.

1. Whisper in a worker, WebGPU first

The worker picks a device and a precision. WebGPU gets fp16 for the encoder when the adapter supports shader-f16, and a 4-bit decoder; the WASM fallback uses 8-bit weights for both:

async function detecterWebGPU() {
  try {
    if (!self.navigator?.gpu) return { ok: false };
    const adaptateur = await navigator.gpu.requestAdapter();
    if (!adaptateur) return { ok: false };
    return { ok: true, fp16: adaptateur.features.has("shader-f16") };
  } catch {
    return { ok: false };
  }
}

const device = gpu.ok ? "webgpu" : "wasm";
const dtype = gpu.ok
  ? { encoder_model: gpu.fp16 ? "fp16" : "fp32", decoder_model_merged: "q4" }
  : { encoder_model: "q8", decoder_model_merged: "q8" };
Enter fullscreen mode Exit fullscreen mode

Two models are offered: onnx-community/whisper-base_timestamped (~90 MB) and whisper-small_timestamped (~250 MB). The _timestamped exports matter: they expose the cross-attention needed for return_timestamps: "word".

navigator.gpu existing is not enough: an adapter can be returned and the pipeline can still fail to build on it. So creation is wrapped and retried in WASM:

try {
  ({ transcripteur: t, device } = await obtenirTranscripteur(modele, false));
} catch (e) {
  self.postMessage({ type: "info", message: `WebGPU indisponible (${e.message || e}), passage en mode compatible.` });
  ({ transcripteur: t, device } = await obtenirTranscripteur(modele, true));
}
Enter fullscreen mode Exit fullscreen mode

There is a second fallback: if word-level timestamps throw, the worker asks for sentence-level timestamps and the UI spreads the words across each sentence, weighted by length.

2. From Whisper chunks to captions

Whisper's word output is messier than it looks. Tokens sometimes arrive split (" micro" + "-entrepreneur."), and noise markers like [Musique] show up. The rule I settled on: a chunk that does not start with a space belongs to the previous word.

const colle = precedent && !/^\s/.test(brut) && !/^[\[(]/.test(texte);
if (colle) {
  precedent.texte += texte;
  precedent.fin = Math.max(precedent.fin, fin);
} else {
  mots.push({ texte, debut, fin });
}
Enter fullscreen mode Exit fullscreen mode

Segmentation is a greedy pass with a few cut conditions: max words per line (user setting, 1 to 5), max 22 characters, a pause longer than 0.6 s, end of sentence, or a comma once the line has two words. Each segment stays on screen 0.3 s after its last word, without ever overlapping the next one. All of this lives in a DOM-free module, so it runs under node --test and also generates the .srt and the .ass karaoke (\k tags) exports.

When the user fixes a typo, the segment keeps its start and end times and the new words are spread inside them. Small, but it means editing never breaks the sync.

3. Rendering: one function, three styles

The same dessinerSousTitres(ctx, segments, t, settings) draws the live preview and the export. It measures each word, wraps at 86% of the 1080 px width, then draws word by word with a thick black stroke, a soft shadow, and a short "pop" on the active word:

const age = t - seg.mots[i].debut;
const pop = estActif ? 1 + 0.1 * Math.max(0, 1 - age / 0.18) : 1;
ctx.save();
ctx.translate(x + w / 2, y);
ctx.scale(pop, pop);
Enter fullscreen mode Exit fullscreen mode

Three styles are just branches in that loop: highlight the active word's color, draw a rounded box behind it, or reveal words only once they've been spoken. The font (Montserrat 900) is self-hosted and preloaded, and the app awaits document.fonts.load() before the first draw. Otherwise the first frames get measured with a fallback font.

Because this module only needs a 2D context, I reused it outside the browser: the promo clip for the tool was rendered in Node with @napi-rs/canvas, the same rendu.js, and transformers.js on CPU.

4. Export: canvas + MediaRecorder

The export plays the video in a hidden <video>, redraws each frame on requestAnimationFrame, and records canvas.captureStream(30) plus the original audio routed through createMediaElementSource:

const candidats = [
  "video/mp4;codecs=avc1.640028,mp4a.40.2",
  "video/mp4;codecs=avc1,mp4a.40.2",
  "video/mp4",
  "video/webm;codecs=vp9,opus",
  "video/webm;codecs=vp8,opus",
  "video/webm",
];
return candidats.find((m) => MediaRecorder.isTypeSupported(m)) || null;
Enter fullscreen mode Exit fullscreen mode

Recent Chrome, Edge and Safari produce MP4; Firefox produces WebM. The honest trade-off: export runs in real time (60 s of video takes about 60 s) and the tab must stay visible, because background tabs throttle rAF. WebCodecs plus a JS muxer would allow faster-than-real-time export and more consistent output. That's the next step, and it isn't shipped yet.

5. The deployment surprise: Hugging Face and *.workers.dev

I first served the app from a *.workers.dev subdomain, and the model never loaded there: model file requests answered 404 whenever the Referer was a workers.dev origin. Same files, same code, different referrer.

Rather than fight it, I moved the static app to Cloudflare Pages (*.pages.dev), where the model loads normally, and left a one-line Worker on the old address that answers 301 to the new one:

const CIBLE = "https://sous-titres.pages.dev";
export default {
  fetch(requete) {
    const url = new URL(requete.url);
    return Response.redirect(CIBLE + url.pathname + url.search, 301);
  },
};
Enter fullscreen mode Exit fullscreen mode

Lesson: if your app fetches third-party assets client-side, test from the real production origin early. The Referer your users send is part of your dependency surface.

Privacy, verifiable

"No upload" is easy to claim, so the FAQ tells users how to check it: open DevTools → Network, run a transcription, and you'll see only the model (Hugging Face), the library (jsDelivr) and anonymous event counters (a visit, an export). No cookies, no third-party analytics.

Limits (the real ones)

  • Desktop Chrome/Edge recommended (that's where I tested it). Firefox and Safari aren't tested yet; without WebGPU the app falls back to WASM, which is slower. Phones can run out of memory.
  • First run downloads the model (~90 MB or ~250 MB), then it's cached.
  • Speed depends on hardware. With WebGPU, a one-minute clip usually transcribes in under a minute. Without it, expect several minutes.
  • Whisper makes mistakes on names, numbers and acronyms. The editor exists for that.
  • Short-form only. Beyond 3–5 minutes, browser memory becomes the bottleneck.
  • Real-time export, as described above.

Try it

sous-titres.pages.dev: free, no account, with a small watermark on exported videos. A one-time paid license removes it. Feedback on the WebGPU path, especially from AMD and Intel GPUs, is very welcome in the comments.

Top comments (0)