DEV Community

Abhishek Pandey
Abhishek Pandey

Posted on AI-assisted

My grandma broke every speech-to-text app, so I fixed the app

Hacktoberfest Weekend Challenge: Build for a Friend Submission ๐Ÿค

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

There is a particular way my siblings and I listen to our grandparents.

We lean in. We catch the shape of a sentence, maybe two thirds of the words, and we nod. Later, in the car, one of us asks our mother what ahay meant, or why the rite happens on the eleventh day and not the tenth, and she tells us, and we lose about half of that too.

I grew up in Mumbai, where Hindi is the water you swim in. It is on the news, in school, in every shop, and it is the Hindi the whole country has agreed on. My grandparents spent their entire lives in a village in Uttar Pradesh, and they speak Awadhi. It sits close enough to Hindi that you feel you ought to follow it, and far enough that you often do not.

So I did the obvious thing. I recorded them and fed it to a transcription app, expecting subtitles for my own family.

What came back was standard Hindi. Not a transcription. A translation, and a lossy one:

what was said what came back both mean
เคนเฅ‹ เคฅเคฏ (ho thay) เคนเฅ‹เคคเคพ เคนเฅˆ (hota hai) "it happens"
เค…เคนเคฏ (ahay) nothing, it vanished "is"
เคชเคนเคฟเคฒเฅ‡ (pahile) เคชเคนเคฒเฅ‡ (pahle) "before"

Every row loses the same thing. The meaning survives and the voice does not. เค…เคนเคฏ is not a typo for เคนเฅˆ. It is how you say "is" where my grandparents are from, and a model trained on newsreaders has simply never met it.

Awadhi has about four million speakers and almost no presence in the data these models learn from, so every one of them treats it as broken Hindi rather than as a language. The correction is not a bug anyone will file. It is the system working exactly as designed, on a dialect nobody designed it for.

Homespun is a speech model retrained to stop doing that.

The friends I built it for are my brother and sister. We are the ones who will want these recordings in twenty years, and we are the ones who cannot fully follow them today.

Demo

Live: homespun-awadhi.onrender.com
Four pages. Listen is the testimony, where every quotation plays in the speaker's own voice with English underneath. Try it runs the retrained model inside your browser, on your own recording, with nothing uploaded. Evidence is the benchmark, and Method covers how it was trained and what that cost.

What changed

Both models read the same held-out recordings through the same script with the same settings. The only difference between them is an 8.7 MB adapter.

before after
Words wrong 0.9323 0.5853
Characters wrong 0.6200 0.3357
Awadhi words kept 5.4% 48.8%

The bottom row is the one I care about. It counts the everyday Awadhi words a general model deletes: เค…เคนเคฏ (is), เคฅเคฏ (the past marker), เคนเฅ‹เคค (happens), เคœเฅŒเคจ (which). You cannot build a sentence without words like these, and the adapter went from keeping one in twenty of them to keeping almost half.

A word error rate above 1.0 is not a typo

The stock model scores worse than 100 percent word error, which sounds impossible until you watch it fail. It does not simply pick wrong words. On a longer clip it gives up and emits one syllable over and over until it runs out of room, so it generates more errors than the reference has words to be wrong about.

said      เคˆ เคฒเคฎเฅเคฌเคพ เคชเฅ‡เคก เค…เคนเคฏ เฅค            "this is a tall tree"
before    เค‡เคœเคผเค‚เคฌเคพ เคชเฅ‡เคฐ เค†เค– เค†เค– เค†เค– เค†เค– เค†เค– เค†เค– เค†เค– เค†เค– เค†เค– เค†เค– เค†เค– ...
after     เคˆ เคฒเคฎเฅเคฌเคพ เคชเฅ‡เคฐ เค…เคนเคฏ เฅค            "this is a tall tree"
Enter fullscreen mode Exit fullscreen mode

The retrained model also decodes about seven times faster, because it stops when the sentence does.

Code

GitHub logo abhishek-pandey7 / Homespun-Voice

Awadhi is treated as broken Hindi by every speech model, so my grandparents' words get "corrected" away. A LoRA adapter on whisper-small that keeps them: 5% to 45% dialect words retained, character error 0.62 to 0.34. 8.7 MB, 52 min on a laptop GPU. Runs in your browser, nothing uploaded.

Homespun

A speech model retrained to write Awadhi down the way it is actually spoken.

Adapter ยท Browser build ยท Benchmark

Why I built this

I grew up in Mumbai. Hindi is everywhere here. It is on the news, in school, in every shop, and it is the Hindi the whole country has agreed on.

My grandparents did not grow up here. They spent their entire lives in a village in Uttar Pradesh, and they speak Awadhi. It sits close enough to Hindi that you feel you ought to follow it, and far enough that you often do not. My siblings and I catch about half of what they say. We nod through the rest, and then one of us asks our parents afterwards.

So I did the obvious thing and pointed a transcription app at a recording. It handed back standard Hindi. Not a transcription, a translation, andโ€ฆ

Both artefacts are published and reusable:

How I Built It

Every model here is open weights. Nothing in the chain needed an API key to produce the result.

The adapter sits on openai/whisper-small, 244M parameters, trained through PEFT with LoRA at rank 16 on the query and value projections, six epochs, over 3 hours 14 minutes of Awadhi speech. The trained weights come to 8.7 MB and took 52 minutes on an RTX 4050 laptop GPU, peaking at 3.7 GB of its 6 GB. That last number is the argument. Teaching a speech model a language it has never heard now costs an evening on hardware you already own.

The corpus is SpeeD-IA (Kumar et al., Interspeech 2022), collected by Dr. Bhimrao Ambedkar University and the Council for Strategic and Defence Research with Karya Inc. and UnReaL-TecE LLP, under CC BY-NC-SA 4.0. It holds 2,538 usable utterances from 18 speakers answering questions about birth, marriage and mourning customs in their own words.

To be exact about provenance: these are not my grandparents' recordings. The motivation is mine, the voices belong to people who gave their time to a research corpus, and the dataset card in the repo says what came from where. Recording my own family is the obvious next step. I wanted the tool to work before asking them to sit in front of a microphone.

The storybook runs on gemma-2-2b-it, quantised to 4-bit NF4, locally. At bf16 it wants roughly 5.2 GB for weights alone and runs out of memory partway through generation on a 6 GB card. NF4 brings it to about 1.5 GB and it fits beside everything else.

Gemma never touches the Awadhi. It returns chapter titles, English narration and a list of utterance ids, and the quotations are looked up from the corpus and inserted afterwards.

That design came out of a failure. My first version asked Gemma to copy each quote and validated afterwards, and it kept truncating them or quietly adding a full stop. One validation pass then reported "3 chapters, 0 quotes, all verbatim" and wrote the file, because every one of its zero quotes was technically verbatim. A check that accepts an empty result is worse than no check, since it reports success. Now the model has no channel through which to alter a word, so fidelity is built into the shape of the pipeline instead of inspected afterwards.

Gemma also translates every quotation into English, which is what makes the site readable to anyone who does not read Devanagari. That needed three guards: reject output that echoes the prompt, reject output that copies a few-shot example, and reject output that repeats a five-word phrase more than twice. The last one caught a translation that looped the same clause 31 times and had already reached the live page.

ElevenLabs (eleven_multilingual_v2) reads the English chapter introductions. The Awadhi is never sent to a speech synthesiser, because the reader plays the speakers' own recordings and synthesising their words would impose Hindi phonology on Awadhi, flattening the thing this project exists to keep. Using the partner's product for the English and deliberately not for the dialect is the whole argument in miniature.

In the browser it runs through Transformers.js against an int8 ONNX build, WebGPU with a WASM fallback, so your audio never leaves your machine. Render hosts the four pages as a static site with no build step, because the model does not run on the server.

Scoring uses jiwer at corpus level rather than an average of per-utterance rates, with normalisation applied identically to both sides. The harness proves itself before it reports anything: a reference scored against itself must give zero, an empty hypothesis must give one, and a single substituted word must match the hand calculation.

Why Does Open Innovation Matter?

Open weights are what made this possible at all.

A closed speech API gives you one lever, which is the audio you send it. There is no prompt that makes a hosted transcription endpoint stop normalising เค…เคนเคฏ into เคนเฅˆ, because the normalising is not a setting. It is what the weights learned. I needed different weights, and the only way to get them was to reach inside.

The economics only work because the model is open. Four million Awadhi speakers is a rounding error to anyone selling transcription, and they are not wrong to skip it. But adapting an open model to a long-tail language now costs one evening and a mid-range laptop GPU, so it does not need funding. It needs one person who cares, and there is at least one of those per language on earth.

The part I did not expect is that measurement needed open weights too. A closed API hands back fluent standard Hindi and looks like it worked. You cannot see what it deleted, because what it deleted is gone. Running the stock model myself over a corpus with human transcriptions is what turned "this feels wrong" into "it keeps 5.4 percent of the dialect markers".

The same argument runs through the rest of the stack. Gemma runs on my machine, so no transcript of my family's speech is uploaded anywhere to be turned into a storybook. Transformers.js means the demo runs in the visitor's browser, so the privacy claim on the site is a fact about the architecture and not a sentence in a policy. For a tool whose whole purpose is keeping family recordings, that is worth the extra work.

What the numbers do not say

The recordings are 8 kHz. Nothing above 4 kHz was ever captured, and that is where much of the energy separating fricatives lives. The ceiling applies to both models equally, so the comparison holds while the absolute figures stay worse than they would be on clean audio.

The same speakers appear in training and testing. The corpus authors split by utterance rather than by speaker, so these results describe adaptation to familiar voices and say nothing about a voice the model has never heard, which is awkwardly the case I care about most.

Long answers improved least. Short prompted sentences gained roughly four times as much as the long spontaneous narratives, which had a fifth of the training data. Doubling the epochs from three to six narrowed that gap without closing it, which points at how much data exists rather than how long it trained.

Twenty of 509 recordings still get worse. The repetition collapse is mostly gone and not entirely; it was 35 before the longer run. The demo page shows one of those failures on purpose, because a reader who discovers a hidden failure stops believing everything else on the page.

Five things that went wrong

data/** in .gitignore silently killed every negation beneath it. Git does not descend into an ignored directory, so !data/raw/manifest.jsonl was dead on arrival. The same trap later 404'd four audio clips in production while they played perfectly on my machine, because locally the files are simply there. Only a request against the deployed URL distinguishes a file that exists from a file that was committed.

PEFT plus reentrant gradient checkpointing severs the autograd graph. With a frozen base, no input requires grad, so backward finds nothing to do and raises. It needs use_reentrant: False.

Batch 8 without gradient checkpointing allocated 9.65 GB on a 6 GB card and did not crash. Windows spills to system RAM rather than raising OOM, so it reported success while running at half the speed of the setting it replaced. Measuring three configurations showed that checkpointing plus a bigger batch beats the obvious alternative on speed and memory together.

Quantising a merged ONNX decoder does nothing at all. It is built around an If node and quantize_dynamic will not descend into subgraphs, so a 739 MB file came out at 739 MB. Quantise the two graphs while they are still flat, then merge.

Blocking repeated n-grams made transcription worse than doing nothing. This one lost twice. It is the obvious fix for a repetition loop, and Awadhi genuinely repeats, so the block deletes real words to prevent a failure affecting 4 percent of clips.

My Agent Session

Curated to the engineering: the corpus traps, the training and quantisation failures, the decode measurements, and the two bugs that were only visible in production. The failures are kept deliberately, since they are the part that shows how the numbers were arrived at.

Homespun: teaching Whisper to hear Awadhi, on a 6 GB laptop GPU
You

My grandparents speak Awadhi and every transcription app turns it into standard Hindi, so my siblings and I lose what they actually said. I want to fine-tune a speech model on the dialect, then build something we can read and listen to. Plan it properly before writing code, and check what this laptop can actually train.

Agent

Probed the machine before planning. RTX 4050 with 6 GB, Python 3.12 present and 3.14 unusable for torch. That rules out hosted notebooks as a necessity: whisper-small LoRA fits in 6 GB.

Wrote a 14-task plan with the labelling bottleneck as the top risk. Hand-correcting 25 minutes of audio is 3 to 4 hours and gates everything downstream, so the plan pushes a walking skeleton through every stage on 3 minutes of audio before that effort is spent.

You

I found an Awadhi speech dataset on GitHub. Check whether it is usable before we commit to recording anything ourselves.

Agent

Checked it rather than assuming. The repo is a 10 KB shell with no audio, but it links SpeeD-IA: an Interspeech 2022 corpus with human transcriptions, an author-defined train/test split, and an explicit CC BY-NC-SA licence.

That removes the labelling bottleneck entirely and makes provenance a citable fact rather than an assumption. Switching to it.

Agent

Three properties of the corpus, found by inspection, none visible from the file listing. Two are traps.

  1. The JSON sidecar beside each clip holds the QUESTION the speaker was asked, not their answer. The spoken response is transcribed only in the TSV. Pairing clips with the sidecar would have trained the model against entirely the wrong words.

  2. Audio is 8 kHz. Everything is upsampled for Whisper, but nothing above 4 kHz was ever captured, and that is where much of the energy separating fricatives lives.

  3. The authors split by utterance, not by speaker, so the same voices appear in train and test. Results describe adaptation to known voices and say nothing about an unheard one.

All three are now in the dataset card and the README.

Agent

Corpus measured: 2,538 usable utterances, 3h14m, 18 speakers. Train 2h36m, test 38m, which lands at 80/20 by duration without anyone choosing it.

Excluded 46: 45 longer than Whisper's 30 second window, 1 under 0.2 seconds. Counted with reasons, not silently dropped.

Also worth recording: data/** in .gitignore silently killed every negation beneath it. Git does not descend into an ignored directory, so !data/raw/manifest.jsonl was dead and the layout would not have survived a clean clone. Verified the fix with git check-ignore in both directions rather than assuming.

You

Before training anything, measure the stock model so we know what we are comparing against.

Agent

Stock whisper-small on the 509-utterance test split: WER 1.0210, CER 0.6335, dialect markers 36 of 709 which is 5.1 percent.

WER above 1.0 is not a typo. The model inserts as well as substitutes: on longer clips it stops transcribing and repeats one syllable until the token limit, so errors outnumber reference words.

The harness verified itself before reporting anything. A reference scored against itself gives 0.0, an empty hypothesis gives 1.0, and one substituted word matches the hand calculation. An unverified harness makes every later number worthless.

Agent

Two bugs before the first real training run.

PEFT plus reentrant gradient checkpointing severs the autograd graph. With a frozen base no input requires grad, so backward finds no grad_fn and raises. It needs use_reentrant: False.

Whisper hard-caps labels at 448 tokens and raises rather than truncating. Three transcripts exceed it, the longest at 1,079. Filtered and counted; truncating would have taught the model to stop mid-sentence.

Agent

Measured three memory configurations instead of guessing.

batch 2 + checkpointing 1.69 GB 1.71 samples/s
batch 4, no checkpointing 5.56 GB 2.47 samples/s
batch 8 + checkpointing 3.65 GB 2.90 samples/s

My assumption that checkpointing was wasted memory was wrong. It frees enough to quadruple the batch, and the larger batch repays the recompute with interest.

The instructive failure: batch 8 WITHOUT checkpointing allocated 9.65 GB on a 6 GB card and did not crash. Windows spills to system RAM rather than raising OOM, so it reported success while running at half the speed of the setting it replaced.

Six epochs, 52 minutes, peak 3.7 GB, adapter 8.7 MB. Validation loss was still falling at three epochs, so the schedule was doubled; every metric improved, including the one I expected to overfit.

You

Now turn the transcripts into something readable. The Awadhi has to stay exactly as the speakers said it, that is the entire point of the project.

Agent

The first design failed, and the failure is the useful part.

I asked Gemma 2B to copy each Awadhi quote verbatim and validated afterwards. At 4-bit it kept truncating quotes or appending a full stop.

Worse, one run reported '3 chapters, 0 quotes, all verbatim' and wrote the file, because every one of its zero quotes was technically verbatim. A check that accepts an empty result is worse than no check, because it reports success.

Rebuilt so the model returns only chapter titles, English narration and a list of utterance ids. The quotations are looked up from the corpus and inserted afterwards. The model now has no channel through which to alter a word, so fidelity is structural rather than verified after the fact.

You

Most people reading this will not read Devanagari. They can see the two transcripts differ but not which one is right, so put the English meaning next to every line.

Agent

Translating the reference for every comparison on the site. Only the reference: a transcript that repeats one syllable forty-six times has no English, and inventing one would flatter it. Those get a computed description instead, which is what a reader actually needs. 'Not a sentence. It repeats เค†เค– 46 times.'

Translation needed three guards, each added after a real failure. The prompt opened with 'It is spoken testimony about family customs' and Gemma returned that exact sentence as the translation for a third of the quotes. Rewritten as few-shot examples, the model copied those examples verbatim instead. The third guard is a repetition check, and it caught a translation looping the same clause 31 times that had already reached the live page.

Long sources degenerate specifically, so a failed attempt now falls back to sentence-by-sentence on the danda, then to word windows for speakers who run twenty seconds without one. 15 of 15 translated, zero loops.

You

A page of numbers is not a demo. People should be able to record or upload their own audio and watch the model work.

Agent

Hugging Face now returns 402 for Gradio Spaces on free CPU, so in-browser instead, which is the better answer anyway. No second service, no cold start, and the audio never leaves the machine. For a tool about keeping family recordings, that is the argument rather than a convenience.

ONNX export hit a real trap. Quantising the merged decoder returned it at 739 MB, exactly its input size, having quantised almost nothing: a merged decoder is built around an If node and quantize_dynamic will not descend into subgraphs. Quantising the two graphs while they are still flat, then merging, gives 186 MB. Total bundle 279 MB, down from 1.76 GB.

Validated the quantised weights through onnxruntime directly, because optimum could not load them: its io-binding helper references torch.int4, which does not exist in torch 2.5. On 20 held-out utterances, int8 costs about 0.04 CER against fp32 and still halves the stock error. Stated on the Method page rather than left to be discovered.

You

Can we get the error rate down further?

Agent

Swept five decode settings on the 40 hardest utterances, where the repetition failure concentrates.

greedy CER 0.5324
beams 5 CER 0.5130
beams 8 CER 0.5161
greedy + no-repeat-4 CER 0.5545
beams 5 + no-repeat-4 CER 0.5404

The n-gram result is the one worth keeping. Blocking repeated 4-grams is the textbook fix for a repetition loop, and it came out worse than greedy, because Awadhi genuinely repeats. The block deletes real words to prevent a failure affecting 4 percent of clips.

Beam 5 applied identically to both models: stock WER 1.0210 to 0.9323, adapted 0.6236 to 0.5853.

You

Go through every clip on the site and tell me which ones the model actually gets right, so I know what to show.

Agent

Ranking them caught a mistake I had shipped an hour earlier.

I had added a repetition penalty to the browser demo to stop it looping. Measured across all 15 published clips, it was the worst of five options, 16 percent worse than doing nothing.

plain greedy CER 0.7208, 1 of 15 looped
penalty 1.25 CER 0.8359, 0 looped <- what I had shipped
penalty 1.1 CER 0.7592, 0 looped
n-gram 3 CER 0.7561, 0 looped
beam 5 CER 0.7177, 0 looped

Beam search removes the loops and is also the most accurate. Both of my additions reverted. That is the second time the intuitive fix for repetition cost more than the repetition did.

You

Here is the deployed URL. Check it from outside before I record anything.

Agent

Deployment had failed first for a reason worth naming: Render found requirements.txt at the repository root, concluded this was a Python app, selected Python 3.14 and tried to install torch 2.5.1 in order to serve static files. No wheel exists for 3.14, the same version trap the README already documents for local setup. Fixed by rooting the service at web/, where no language is detected.

Checking the live URL then found something local testing cannot catch: four example clips 404ing in production. Their .wav files were still caught by the blanket audio ignore. Locally the files are simply there, so only a request against the deployed URL distinguishes a file that exists from a file that was committed. Those four are also the fallback when a visitor has no WebGPU, so they were the one thing that had to work.

Final numbers. WER 0.9323 to 0.5853. CER 0.6200 to 0.3357. Awadhi marker words kept 5.4 percent to 48.8 percent. Adapter 8.7 MB, 52 minutes on a laptop GPU.

Prize Categories

Best Use of Gemma. gemma-2-2b-it at 4-bit, running locally, assembling the storybook and translating every quotation, architecturally prevented from altering the Awadhi it is arranging.

Best Use of Render. Hosting the four-page reader, including the in-browser model demo.

Best Use of ElevenLabs. eleven_multilingual_v2 narrating the English chapters, deliberately not the dialect.

Top comments (0)