Training a wake word that actually fires
I've been building an always-listening voice assistant that runs entirely on a Raspberry Pi 4.
The very first thing it has to do is also the easiest to underestimate: notice when you say its
name. If the wake word doesn't fire, nothing else in the pipeline ever gets a turn — the
speech-to-text, the LLM, the voice, all of it sits there waiting.
My wake word is "Nova". Getting it to trigger reliably took me down a rabbit hole, and most of what
actually moved the needle was not where I expected. This post is the recipe I wish I'd had — with
the specific mistakes that cost me recall, so you can skip them.
Everything here uses open-source tools: openWakeWord
for the model and Piper for synthetic speech. The approach
works for any keyword, not just mine.
What a wake word model actually is
openWakeWord doesn't train a giant speech model. It trains a small classifier on top of a
shared, pre-trained audio embedding. The pipeline is:
audio → melspectrogram → shared embedding → small "is this the wake word?" classifier
That's why you can train it on a laptop or a free GPU notebook and run it on a 2 GB Pi: only the
tiny classifier is yours. You feed it two things: positives (lots of people saying your word)
and negatives (speech and noise that is not your word). The catch is that you rarely have
thousands of real recordings of a made-up name — so you synthesize them.
Step 1 — Synthesize positives with TTS (and the lesson that cost me the most)
The standard trick is to generate thousands of utterances of your wake word with a text-to-speech
engine, varying voice, speed and pitch. I used Piper.
Here is the lesson, and it's the big one: the language of the TTS voice has to match how you'll
actually say the word. My first model was trained with English voices. In English, "Nova" is
said roughly NOH-vah; in Spanish (how I say it) it's /ˈno.βa/. The model dutifully learned the
English pronunciation — and then barely fired when I spoke to it. The utterance would score around
0.002. Not "a bit low". Essentially zero.
Switching the positives to Spanish Piper voices was the single biggest improvement I made.
After listening to a bunch of candidates, I kept the three that pronounced "Nova" cleanly
(es_ES-sharvard-medium, es_MX-ald-medium, es_MX-claude-high) and dropped the ones that sounded
off on this particular word.
# Grab a few Spanish voices and generate a small batch to listen to FIRST
python -m piper.download_voices --download-dir voices \
es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
# then synthesize many "nova" clips at 16 kHz mono, varying voice/speed/pitch
Do this before you train anything: generate ~50 clips and actually listen. One minute of
listening told me the English voices were wrong before I'd burned a single GPU-hour. Whatever you
hear in those clips is what the model is about to learn.
Step 2 — Negatives and augmentation
Positives alone teach the model to say "yes"; it also has to learn to say "no". openWakeWord's
training uses large public datasets of general speech and noise as negatives, plus pre-computed
features to validate false positives. On top of that it augments the positives by mixing in
noise and room impulse responses (RIRs) so the model survives a real room instead of only clean TTS.
You mostly get this for free from the project's training notebook — just don't skip it.
Step 3 — Train on a free GPU notebook (watch the hardware)
Training downloads several GB of audio and runs for a while, so I don't do it on the Pi. Google
Colab and Kaggle both work and are free. The config change that makes it your word is a single
line:
target_phrase = ["nova"]
Two hardware gotchas that wasted my time:
- On Kaggle, pick the T4 GPU, not the P100. The P100 is Pascal-era and Kaggle's bundled PyTorch won't run on it. Kaggle also has a killer feature for this: Save & Run All (Commit) executes the whole notebook on their servers, to completion (up to 12 h), with your browser closed.
- On Colab (free), the session dies on inactivity and won't run in the background — you have to keep the tab open, and pin the runtime to the Python version the notebook expects.
Out comes a single nova.onnx file. That's the whole model.
Step 4 — The integration bug that made it score zero
With the model in place, it still wouldn't fire — and this one had nothing to do with training.
openWakeWord expects to be fed audio in 1280-sample windows (80 ms at 16 kHz). My audio capture
was handing it 480-sample frames (30 ms), because that's the frame size the voice-activity
detector wanted. Fed the wrong window size, the detector returned scores near 0 on every frame
and never triggered. The fix was to buffer incoming frames up to 1280 samples before calling
predict():
# accumulate frames until we have a full 1280-sample window, then score
buffer.extend(frame)
while len(buffer) >= 1280:
window, buffer = buffer[:1280], buffer[1280:]
score = model.predict(window)
If your freshly trained model scores zero on everything, suspect the plumbing before you blame the
training.
Step 5 — Measure recall, don't trust your ears
Early on I "tested" the wake word by saying it a few times and nodding. That's how you fool
yourself. I built a tiny evaluation harness instead: a folder of positive clips, a folder of
negatives, and a sweep across thresholds that prints recall and false positives per threshold.
evaluate_wakeword --positives eval/positives --negatives eval/negatives \
--thresholds 0.2,0.3,0.35,0.4,0.5
Use at least ~30 varied clips — different distances, speeds, background noise. In my set, about
a dozen of them were specifically the kind that fool a naive model, and they're the ones that tell
you the truth.
This also killed a tempting assumption: that the detection threshold is the lever for recall.
It isn't. When the pronunciation genuinely doesn't match, the utterance scores ~0.002, and no
threshold saves you — in fact lowering it from 0.4 to 0.3 made things worse (more false
positives, no real gain). The threshold is a fine-tuning dial, not a fix for bad training data.
Step 6 — The real jump: fine-tune with your own voice
The synthetic-only model learned a TTS "Nova", not my voice in my room through my microphone.
The biggest, most reliable win was adding real recordings of me saying the word and retraining.
I wrote a small recorder that beeps and captures clips at 16 kHz mono, then varied distance, tone,
speed and background noise across ~60 of them (and held ~10% back for evaluation). In the training
notebook, those real positives get upsampled and mixed in with the synthetic ones.
The result, measured on the same evaluation set:
Recall went from 53% → 70% at threshold 0.35, just by folding in real-voice positives.
Not magic, but a real, measured step up — and exactly the kind of improvement that's invisible if
you're only judging by ear.
Reproduce it yourself (the short version)
Here's the minimal path with public tools, so you can train your own keyword. Swap "nova" for
yours throughout.
1. Install and synthesize positives. Generate a few thousand clips, varying voice and speed:
pip install piper-tts openwakeword
python -m piper.download_voices --download-dir voices \
es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
import random, subprocess, pathlib
pathlib.Path("positives").mkdir(exist_ok=True)
voices = ["es_ES-sharvard-medium", "es_MX-ald-medium", "es_MX-claude-high"]
for i in range(2000):
v = random.choice(voices)
length = round(random.uniform(0.9, 1.3), 2) # speed variation
subprocess.run(
["piper", "--model", f"voices/{v}.onnx",
"--length-scale", str(length), "--output_file", f"positives/{v}_{i}.wav"],
input=b"nova\n",
)
# (CLI flags vary slightly by Piper version — check `piper --help`.)
Then listen to a handful before going further.
2. Train. Open openWakeWord's automatic_model_training.ipynb
(in their repo) on Kaggle or Colab, select a T4
GPU, set target_phrase = ["nova"], point it at your positives/ folder, and Run All. It pulls
the negative/background datasets and augmentation for you and exports a single nova.onnx.
3. Evaluate against your own clip folders and sweep thresholds — this is the step people skip:
import os, numpy as np, soundfile as sf
from openwakeword.model import Model
model = Model(wakeword_models=["nova.onnx"])
def best_score(path):
model.reset()
audio, _ = sf.read(path) # 16 kHz mono
audio = (audio * 32767).astype(np.int16)
top = 0.0
for i in range(0, len(audio) - 1280, 1280): # 80 ms windows
top = max(top, model.predict(audio[i:i + 1280])["nova"])
return top
pos = [best_score(f"eval/positives/{f}") for f in os.listdir("eval/positives")]
neg = [best_score(f"eval/negatives/{f}") for f in os.listdir("eval/negatives")]
for t in (0.2, 0.3, 0.35, 0.4, 0.5):
recall = sum(s >= t for s in pos) / len(pos)
fp = sum(s >= t for s in neg)
print(f"thr={t}: recall={recall:.0%} false_positives={fp}/{len(neg)}")
4. Fine-tune with real voice (optional, high impact). Record ~60 clips of yourself saying the
word (16 kHz mono — a few lines with sounddevice), hold back ~10% for evaluation, and add them to
the training set with upsampling. This is what took me from 53% to 70%.
5. Deploy. Copy nova.onnx to your device and feed the detector 1280-sample windows (see
the buffering snippet above). Tune the threshold from your evaluation numbers.
What actually moved the needle
If you only remember four things:
- Match the TTS language to how you'll say the word. This was worth more than any hyperparameter.
- Listen to your synthetic positives before training. One minute saves hours.
- Measure recall and false positives on a varied clip set. Your ears lie; a threshold sweep doesn't.
- Fine-tune with your own real voice. Synthetic gets you started; real recordings get you reliable.
And one bonus, because it bit me hardest: if a model scores zero on everything, check the window
size you're feeding it before you retrain anything.
Have you trained a custom wake word? I'd love to hear what moved recall for you — especially for
non-English keywords.
Want more context, or to see how we did it?
Nova isn't fully public yet, but you can get early access to the repository — all the
documentation and the complete source code — at https://gitlab.com/gabrielhruiz1/nova.
Leave us a message and we'll try to grant you access ASAP, until we publish everything
officially (we're still working on a few parts).
Top comments (0)