This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
There is a particular way my siblings and I listen to our grandparents.
We lean in. We catch the shape of a sentence, maybe two thirds of the words, and we nod. Later, in the car, one of us asks our mother what ahay meant, or why the rite happens on the eleventh day and not the tenth, and she tells us, and we lose about half of that too.
I grew up in Mumbai, where Hindi is the water you swim in. It is on the news, in school, in every shop, and it is the Hindi the whole country has agreed on. My grandparents spent their entire lives in a village in Uttar Pradesh, and they speak Awadhi. It sits close enough to Hindi that you feel you ought to follow it, and far enough that you often do not.
So I did the obvious thing. I recorded them and fed it to a transcription app, expecting subtitles for my own family.
What came back was standard Hindi. Not a transcription. A translation, and a lossy one:
| what was said | what came back | both mean |
|---|---|---|
เคนเฅ เคฅเคฏ (ho thay)
|
เคนเฅเคคเคพ เคนเฅ (hota hai)
|
"it happens" |
เค
เคนเคฏ (ahay)
|
nothing, it vanished | "is" |
เคชเคนเคฟเคฒเฅ (pahile)
|
เคชเคนเคฒเฅ (pahle)
|
"before" |
Every row loses the same thing. The meaning survives and the voice does not. เค
เคนเคฏ is not a typo for เคนเฅ. It is how you say "is" where my grandparents are from, and a model trained on newsreaders has simply never met it.
Awadhi has about four million speakers and almost no presence in the data these models learn from, so every one of them treats it as broken Hindi rather than as a language. The correction is not a bug anyone will file. It is the system working exactly as designed, on a dialect nobody designed it for.
Homespun is a speech model retrained to stop doing that.
The friends I built it for are my brother and sister. We are the ones who will want these recordings in twenty years, and we are the ones who cannot fully follow them today.
Demo
Live: homespun-awadhi.onrender.com
Four pages. Listen is the testimony, where every quotation plays in the speaker's own voice with English underneath. Try it runs the retrained model inside your browser, on your own recording, with nothing uploaded. Evidence is the benchmark, and Method covers how it was trained and what that cost.
What changed
Both models read the same held-out recordings through the same script with the same settings. The only difference between them is an 8.7 MB adapter.
| before | after | |
|---|---|---|
| Words wrong | 0.9323 | 0.5853 |
| Characters wrong | 0.6200 | 0.3357 |
| Awadhi words kept | 5.4% | 48.8% |
The bottom row is the one I care about. It counts the everyday Awadhi words a general model deletes: เค
เคนเคฏ (is), เคฅเคฏ (the past marker), เคนเฅเคค (happens), เคเฅเคจ (which). You cannot build a sentence without words like these, and the adapter went from keeping one in twenty of them to keeping almost half.
A word error rate above 1.0 is not a typo
The stock model scores worse than 100 percent word error, which sounds impossible until you watch it fail. It does not simply pick wrong words. On a longer clip it gives up and emits one syllable over and over until it runs out of room, so it generates more errors than the reference has words to be wrong about.
said เค เคฒเคฎเฅเคฌเคพ เคชเฅเคก เค
เคนเคฏ เฅค "this is a tall tree"
before เคเคเคผเคเคฌเคพ เคชเฅเคฐ เคเค เคเค เคเค เคเค เคเค เคเค เคเค เคเค เคเค เคเค เคเค ...
after เค เคฒเคฎเฅเคฌเคพ เคชเฅเคฐ เค
เคนเคฏ เฅค "this is a tall tree"
The retrained model also decodes about seven times faster, because it stops when the sentence does.
Code
abhishek-pandey7
/
Homespun-Voice
Awadhi is treated as broken Hindi by every speech model, so my grandparents' words get "corrected" away. A LoRA adapter on whisper-small that keeps them: 5% to 45% dialect words retained, character error 0.62 to 0.34. 8.7 MB, 52 min on a laptop GPU. Runs in your browser, nothing uploaded.
Homespun
A speech model retrained to write Awadhi down the way it is actually spoken.
Adapter ยท Browser build ยท Benchmark
Why I built this
I grew up in Mumbai. Hindi is everywhere here. It is on the news, in school, in every shop, and it is the Hindi the whole country has agreed on.
My grandparents did not grow up here. They spent their entire lives in a village in Uttar Pradesh, and they speak Awadhi. It sits close enough to Hindi that you feel you ought to follow it, and far enough that you often do not. My siblings and I catch about half of what they say. We nod through the rest, and then one of us asks our parents afterwards.
So I did the obvious thing and pointed a transcription app at a recording. It handed back standard Hindi. Not a transcription, a translation, andโฆ
Both artefacts are published and reusable:
- abhshkp/homespun-awadhi-lora, the adapter, 8.7 MB
- abhshkp/homespun-awadhi-web, the int8 ONNX build for the browser, 279 MB
How I Built It
Every model here is open weights. Nothing in the chain needed an API key to produce the result.
The adapter sits on openai/whisper-small, 244M parameters, trained through PEFT with LoRA at rank 16 on the query and value projections, six epochs, over 3 hours 14 minutes of Awadhi speech. The trained weights come to 8.7 MB and took 52 minutes on an RTX 4050 laptop GPU, peaking at 3.7 GB of its 6 GB. That last number is the argument. Teaching a speech model a language it has never heard now costs an evening on hardware you already own.
The corpus is SpeeD-IA (Kumar et al., Interspeech 2022), collected by Dr. Bhimrao Ambedkar University and the Council for Strategic and Defence Research with Karya Inc. and UnReaL-TecE LLP, under CC BY-NC-SA 4.0. It holds 2,538 usable utterances from 18 speakers answering questions about birth, marriage and mourning customs in their own words.
To be exact about provenance: these are not my grandparents' recordings. The motivation is mine, the voices belong to people who gave their time to a research corpus, and the dataset card in the repo says what came from where. Recording my own family is the obvious next step. I wanted the tool to work before asking them to sit in front of a microphone.
The storybook runs on gemma-2-2b-it, quantised to 4-bit NF4, locally. At bf16 it wants roughly 5.2 GB for weights alone and runs out of memory partway through generation on a 6 GB card. NF4 brings it to about 1.5 GB and it fits beside everything else.
Gemma never touches the Awadhi. It returns chapter titles, English narration and a list of utterance ids, and the quotations are looked up from the corpus and inserted afterwards.
That design came out of a failure. My first version asked Gemma to copy each quote and validated afterwards, and it kept truncating them or quietly adding a full stop. One validation pass then reported "3 chapters, 0 quotes, all verbatim" and wrote the file, because every one of its zero quotes was technically verbatim. A check that accepts an empty result is worse than no check, since it reports success. Now the model has no channel through which to alter a word, so fidelity is built into the shape of the pipeline instead of inspected afterwards.
Gemma also translates every quotation into English, which is what makes the site readable to anyone who does not read Devanagari. That needed three guards: reject output that echoes the prompt, reject output that copies a few-shot example, and reject output that repeats a five-word phrase more than twice. The last one caught a translation that looped the same clause 31 times and had already reached the live page.
ElevenLabs (eleven_multilingual_v2) reads the English chapter introductions. The Awadhi is never sent to a speech synthesiser, because the reader plays the speakers' own recordings and synthesising their words would impose Hindi phonology on Awadhi, flattening the thing this project exists to keep. Using the partner's product for the English and deliberately not for the dialect is the whole argument in miniature.
In the browser it runs through Transformers.js against an int8 ONNX build, WebGPU with a WASM fallback, so your audio never leaves your machine. Render hosts the four pages as a static site with no build step, because the model does not run on the server.
Scoring uses jiwer at corpus level rather than an average of per-utterance rates, with normalisation applied identically to both sides. The harness proves itself before it reports anything: a reference scored against itself must give zero, an empty hypothesis must give one, and a single substituted word must match the hand calculation.
Why Does Open Innovation Matter?
Open weights are what made this possible at all.
A closed speech API gives you one lever, which is the audio you send it. There is no prompt that makes a hosted transcription endpoint stop normalising เค
เคนเคฏ into เคนเฅ, because the normalising is not a setting. It is what the weights learned. I needed different weights, and the only way to get them was to reach inside.
The economics only work because the model is open. Four million Awadhi speakers is a rounding error to anyone selling transcription, and they are not wrong to skip it. But adapting an open model to a long-tail language now costs one evening and a mid-range laptop GPU, so it does not need funding. It needs one person who cares, and there is at least one of those per language on earth.
The part I did not expect is that measurement needed open weights too. A closed API hands back fluent standard Hindi and looks like it worked. You cannot see what it deleted, because what it deleted is gone. Running the stock model myself over a corpus with human transcriptions is what turned "this feels wrong" into "it keeps 5.4 percent of the dialect markers".
The same argument runs through the rest of the stack. Gemma runs on my machine, so no transcript of my family's speech is uploaded anywhere to be turned into a storybook. Transformers.js means the demo runs in the visitor's browser, so the privacy claim on the site is a fact about the architecture and not a sentence in a policy. For a tool whose whole purpose is keeping family recordings, that is worth the extra work.
What the numbers do not say
The recordings are 8 kHz. Nothing above 4 kHz was ever captured, and that is where much of the energy separating fricatives lives. The ceiling applies to both models equally, so the comparison holds while the absolute figures stay worse than they would be on clean audio.
The same speakers appear in training and testing. The corpus authors split by utterance rather than by speaker, so these results describe adaptation to familiar voices and say nothing about a voice the model has never heard, which is awkwardly the case I care about most.
Long answers improved least. Short prompted sentences gained roughly four times as much as the long spontaneous narratives, which had a fifth of the training data. Doubling the epochs from three to six narrowed that gap without closing it, which points at how much data exists rather than how long it trained.
Twenty of 509 recordings still get worse. The repetition collapse is mostly gone and not entirely; it was 35 before the longer run. The demo page shows one of those failures on purpose, because a reader who discovers a hidden failure stops believing everything else on the page.
Five things that went wrong
data/** in .gitignore silently killed every negation beneath it. Git does not descend into an ignored directory, so !data/raw/manifest.jsonl was dead on arrival. The same trap later 404'd four audio clips in production while they played perfectly on my machine, because locally the files are simply there. Only a request against the deployed URL distinguishes a file that exists from a file that was committed.
PEFT plus reentrant gradient checkpointing severs the autograd graph. With a frozen base, no input requires grad, so backward finds nothing to do and raises. It needs use_reentrant: False.
Batch 8 without gradient checkpointing allocated 9.65 GB on a 6 GB card and did not crash. Windows spills to system RAM rather than raising OOM, so it reported success while running at half the speed of the setting it replaced. Measuring three configurations showed that checkpointing plus a bigger batch beats the obvious alternative on speed and memory together.
Quantising a merged ONNX decoder does nothing at all. It is built around an If node and quantize_dynamic will not descend into subgraphs, so a 739 MB file came out at 739 MB. Quantise the two graphs while they are still flat, then merge.
Blocking repeated n-grams made transcription worse than doing nothing. This one lost twice. It is the obvious fix for a repetition loop, and Awadhi genuinely repeats, so the block deletes real words to prevent a failure affecting 4 percent of clips.
My Agent Session
Curated to the engineering: the corpus traps, the training and quantisation failures, the decode measurements, and the two bugs that were only visible in production. The failures are kept deliberately, since they are the part that shows how the numbers were arrived at.
Prize Categories
Best Use of Gemma. gemma-2-2b-it at 4-bit, running locally, assembling the storybook and translating every quotation, architecturally prevented from altering the Awadhi it is arranging.
Best Use of Render. Hosting the four-page reader, including the in-browser model demo.
Best Use of ElevenLabs. eleven_multilingual_v2 narrating the English chapters, deliberately not the dialect.
Top comments (0)