## What I Built
I built Saaz, a subtitle tool for regional-language video. You upload a clip, and it returns timed Hindi subtitles plus an English translation. You can edit them in the browser and export
.srt or .vtt. Everything runs on your own machine with open-weight models, and no closed AI API is called at any point.
I built it for Raashi, who edit reels they caption footage for clients, and uploading unreleased footage to a hosted captioning service wasn't an option.
That is the problem Saaz solves. When you caption someone's unreleased footage, the video, the transcript and the translation normally all end up on somebody else's server. With Saaz every byte stays local, and the offline claim is checkable: npm run prove:offline runs the pipeline with all outbound network poisoned.
It also doesn't pretend to be perfect. A translation often needs more characters than the on-screen time window allows, so a Fit stage escalates through five outcomes:
- verbatim: it already fits
- reflowed: a small model rewrites it inside the budget
- extended: borrow time from the following gap
- split: divide the cue deterministically and retranslate both halves
- refused: nothing works, so flag it for a human and say why
The last outcome matters most. A subtitle file is something a client signs off on, so a tool that always outputs something is worse than one that tells you which lines need a person.
Code
https://github.com/Virtualord/saaz
How I Built It
The pipeline is: probe → decode → Whisper ASR → Silero VAD → align → segment → glossary → OPUS-MT → Fit.
| Stage | Model | Licence |
|---|---|---|
| Speech recognition | onnx-community/whisper-small |
MIT |
| Speech boundaries | onnx-community/silero-vad |
MIT |
| Translation draft | Xenova/opus-mt-hi-en |
Apache-2.0 |
| Caption rewrite | HuggingFaceTB/SmolLM2-360M-Instruct |
Apache-2.0 |
All four run through one runtime, transformers.js over onnxruntime-node, quantised to q8 so they fit in CPU RAM with no GPU. The app is TypeScript throughout, with a React + shadcn/ui + Tailwind v4 frontend, Vite, SQLite for jobs, and Vitest for the unit tests (segmenter, fit, align, glossary, export). I deploy it from a Docker image to Render, with the weights baked in at build time so the live service never contacts Hugging Face.
I chose every default from a benchmark rather than reputation:
-
Whisper-base emits Arabic script for Hindi audio (CER 1.00). whisper-small scores 0.165. I also tried
whisper-large-v3-turbo, but atq8it returned empty transcripts, so I removed it instead of shipping it broken. - Qwen2.5-0.5B and 1.5B couldn't do the caption task. Both failed to produce parseable JSON and degenerated into repetition. SmolLM2-360M returned valid JSON first try and ran 2.5× faster. Smaller model, better result.
- The Whisper export has no cross-attention outputs, so it can't emit word timestamps. Its segment timings collapsed into uniform 3-second blocks that matched nothing in the audio. I found this by inspecting real output and fixed it with Silero VAD.
-
Open MT has vocabulary gaps.
मालाई(the sweet) translated to "by Miley". I ruled out tokenisation bugs, so Saaz ships a user-pinned glossary that is substituted before translation.
There is deliberately no agent loop in the fit path. A subtitle deliverable has to be reproducible, so escalation is deterministic.
Honest limitations: Whisper quality on regional languages is imperfect (hence the editor), word timing is approximated by proportional alignment, and CPU inference runs at roughly 4× realtime for ASR. That is the price of keeping footage on a laptop.
Why Does Open Innovation Matter?
A closed API would have made this project impossible, because the whole point is that the footage never leaves the machine. Open weights let me run speech recognition, translation and caption rewriting in-process and prove it with the network unplugged.
Openness also made the work inspectable. I could open the Whisper export and discover it had no word timestamps. I could swap Qwen out for SmolLM2 when Qwen failed, and I could read every model's licence and put it in the UI. I even left Gemma out of the defaults because its licence isn't OSI-approved. With a closed API I would have had to trust whatever it did to my friend's footage. Here I can show exactly what ran and why.
Prize Categories
-
Best Use of Render: Saaz ships a
render.yamlblueprint and deploys from a Docker image with weights baked in.
Top comments (0)