<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gabriel Hidalgo</title>
    <description>The latest articles on DEV Community by Gabriel Hidalgo (@gabrielhruiz).</description>
    <link>https://dev.to/gabrielhruiz</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4170994%2F32072aa7-d272-4963-b307-c1a316fd6781.jpg</url>
      <title>DEV Community: Gabriel Hidalgo</title>
      <link>https://dev.to/gabrielhruiz</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gabrielhruiz"/>
    <language>en</language>
    <item>
      <title>Training a wake word that actually fires: lessons from an offline voice assistant</title>
      <dc:creator>Gabriel Hidalgo</dc:creator>
      <pubDate>Thu, 08 Oct 2026 11:14:26 +0000</pubDate>
      <link>https://dev.to/gabrielhruiz/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant-3716</link>
      <guid>https://dev.to/gabrielhruiz/training-a-wake-word-that-actually-fires-lessons-from-an-offline-voice-assistant-3716</guid>
      <description>&lt;h1&gt;
  
  
  Training a wake word that actually fires
&lt;/h1&gt;

&lt;p&gt;I've been building an always-listening voice assistant that runs entirely on a Raspberry Pi 4.&lt;br&gt;
The very first thing it has to do is also the easiest to underestimate: notice when you say its&lt;br&gt;
name. If the &lt;strong&gt;wake word&lt;/strong&gt; doesn't fire, nothing else in the pipeline ever gets a turn — the&lt;br&gt;
speech-to-text, the LLM, the voice, all of it sits there waiting.&lt;/p&gt;

&lt;p&gt;My wake word is "Nova". Getting it to trigger reliably took me down a rabbit hole, and most of what&lt;br&gt;
actually moved the needle was &lt;em&gt;not&lt;/em&gt; where I expected. This post is the recipe I wish I'd had — with&lt;br&gt;
the specific mistakes that cost me recall, so you can skip them.&lt;/p&gt;

&lt;p&gt;Everything here uses open-source tools: &lt;a href="https://github.com/dscripka/openWakeWord" rel="noopener noreferrer"&gt;openWakeWord&lt;/a&gt;&lt;br&gt;
for the model and &lt;a href="https://github.com/rhasspy/piper" rel="noopener noreferrer"&gt;Piper&lt;/a&gt; for synthetic speech. The approach&lt;br&gt;
works for any keyword, not just mine.&lt;/p&gt;
&lt;h2&gt;
  
  
  What a wake word model actually is
&lt;/h2&gt;

&lt;p&gt;openWakeWord doesn't train a giant speech model. It trains a &lt;strong&gt;small classifier&lt;/strong&gt; on top of a&lt;br&gt;
shared, pre-trained audio embedding. The pipeline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;audio → melspectrogram → shared embedding → small "is this the wake word?" classifier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's why you can train it on a laptop or a free GPU notebook and run it on a 2 GB Pi: only the&lt;br&gt;
tiny classifier is yours. You feed it two things: &lt;strong&gt;positives&lt;/strong&gt; (lots of people saying your word)&lt;br&gt;
and &lt;strong&gt;negatives&lt;/strong&gt; (speech and noise that is &lt;em&gt;not&lt;/em&gt; your word). The catch is that you rarely have&lt;br&gt;
thousands of real recordings of a made-up name — so you synthesize them.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 1 — Synthesize positives with TTS (and the lesson that cost me the most)
&lt;/h2&gt;

&lt;p&gt;The standard trick is to generate thousands of utterances of your wake word with a text-to-speech&lt;br&gt;
engine, varying voice, speed and pitch. I used Piper.&lt;/p&gt;

&lt;p&gt;Here is the lesson, and it's the big one: &lt;strong&gt;the language of the TTS voice has to match how you'll&lt;br&gt;
actually say the word.&lt;/strong&gt; My first model was trained with English voices. In English, "Nova" is&lt;br&gt;
said roughly &lt;em&gt;NOH-vah&lt;/em&gt;; in Spanish (how I say it) it's &lt;em&gt;/ˈno.βa/&lt;/em&gt;. The model dutifully learned the&lt;br&gt;
English pronunciation — and then barely fired when I spoke to it. The utterance would score around&lt;br&gt;
&lt;strong&gt;0.002&lt;/strong&gt;. Not "a bit low". Essentially zero.&lt;/p&gt;

&lt;p&gt;Switching the positives to &lt;strong&gt;Spanish Piper voices&lt;/strong&gt; was the single biggest improvement I made.&lt;br&gt;
After listening to a bunch of candidates, I kept the three that pronounced "Nova" cleanly&lt;br&gt;
(&lt;code&gt;es_ES-sharvard-medium&lt;/code&gt;, &lt;code&gt;es_MX-ald-medium&lt;/code&gt;, &lt;code&gt;es_MX-claude-high&lt;/code&gt;) and dropped the ones that sounded&lt;br&gt;
off on this particular word.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Grab a few Spanish voices and generate a small batch to listen to FIRST&lt;/span&gt;
python &lt;span class="nt"&gt;-m&lt;/span&gt; piper.download_voices &lt;span class="nt"&gt;--download-dir&lt;/span&gt; voices &lt;span class="se"&gt;\&lt;/span&gt;
    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
&lt;span class="c"&gt;# then synthesize many "nova" clips at 16 kHz mono, varying voice/speed/pitch&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do this before you train anything:&lt;/strong&gt; generate ~50 clips and actually &lt;em&gt;listen&lt;/em&gt;. One minute of&lt;br&gt;
listening told me the English voices were wrong before I'd burned a single GPU-hour. Whatever you&lt;br&gt;
hear in those clips is what the model is about to learn.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Step 2 — Negatives and augmentation
&lt;/h2&gt;

&lt;p&gt;Positives alone teach the model to say "yes"; it also has to learn to say "no". openWakeWord's&lt;br&gt;
training uses large public datasets of general speech and noise as negatives, plus pre-computed&lt;br&gt;
features to validate false positives. On top of that it &lt;strong&gt;augments&lt;/strong&gt; the positives by mixing in&lt;br&gt;
noise and room impulse responses (RIRs) so the model survives a real room instead of only clean TTS.&lt;br&gt;
You mostly get this for free from the project's training notebook — just don't skip it.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 3 — Train on a free GPU notebook (watch the hardware)
&lt;/h2&gt;

&lt;p&gt;Training downloads several GB of audio and runs for a while, so I don't do it on the Pi. Google&lt;br&gt;
Colab and Kaggle both work and are free. The config change that makes it &lt;em&gt;your&lt;/em&gt; word is a single&lt;br&gt;
line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;target_phrase&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nova&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two hardware gotchas that wasted my time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;On Kaggle, pick the T4 GPU, not the P100.&lt;/strong&gt; The P100 is Pascal-era and Kaggle's bundled PyTorch
won't run on it. Kaggle also has a killer feature for this: &lt;em&gt;Save &amp;amp; Run All (Commit)&lt;/em&gt; executes the
whole notebook on their servers, to completion (up to 12 h), with your browser closed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On Colab (free), the session dies on inactivity&lt;/strong&gt; and won't run in the background — you have to
keep the tab open, and pin the runtime to the Python version the notebook expects.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Out comes a single &lt;code&gt;nova.onnx&lt;/code&gt; file. That's the whole model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4 — The integration bug that made it score zero
&lt;/h2&gt;

&lt;p&gt;With the model in place, it still wouldn't fire — and this one had nothing to do with training.&lt;/p&gt;

&lt;p&gt;openWakeWord expects to be fed audio in &lt;strong&gt;1280-sample windows (80 ms at 16 kHz)&lt;/strong&gt;. My audio capture&lt;br&gt;
was handing it &lt;strong&gt;480-sample frames (30 ms)&lt;/strong&gt;, because that's the frame size the voice-activity&lt;br&gt;
detector wanted. Fed the wrong window size, the detector returned scores near &lt;strong&gt;0&lt;/strong&gt; on every frame&lt;br&gt;
and never triggered. The fix was to buffer incoming frames up to 1280 samples before calling&lt;br&gt;
&lt;code&gt;predict()&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# accumulate frames until we have a full 1280-sample window, then score
&lt;/span&gt;&lt;span class="nb"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;extend&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;1280&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;window&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;buffer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;[:&lt;/span&gt;&lt;span class="mi"&gt;1280&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nb"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1280&lt;/span&gt;&lt;span class="p"&gt;:]&lt;/span&gt;
    &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;window&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your freshly trained model scores zero on &lt;em&gt;everything&lt;/em&gt;, suspect the plumbing before you blame the&lt;br&gt;
training.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 5 — Measure recall, don't trust your ears
&lt;/h2&gt;

&lt;p&gt;Early on I "tested" the wake word by saying it a few times and nodding. That's how you fool&lt;br&gt;
yourself. I built a tiny evaluation harness instead: a folder of positive clips, a folder of&lt;br&gt;
negatives, and a sweep across thresholds that prints recall and false positives per threshold.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;evaluate_wakeword &lt;span class="nt"&gt;--positives&lt;/span&gt; &lt;span class="nb"&gt;eval&lt;/span&gt;/positives &lt;span class="nt"&gt;--negatives&lt;/span&gt; &lt;span class="nb"&gt;eval&lt;/span&gt;/negatives &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--thresholds&lt;/span&gt; 0.2,0.3,0.35,0.4,0.5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use &lt;strong&gt;at least ~30 varied clips&lt;/strong&gt; — different distances, speeds, background noise. In my set, about&lt;br&gt;
a dozen of them were specifically the kind that fool a naive model, and they're the ones that tell&lt;br&gt;
you the truth.&lt;/p&gt;

&lt;p&gt;This also killed a tempting assumption: that the &lt;strong&gt;detection threshold&lt;/strong&gt; is the lever for recall.&lt;br&gt;
It isn't. When the pronunciation genuinely doesn't match, the utterance scores ~0.002, and no&lt;br&gt;
threshold saves you — in fact lowering it from 0.4 to 0.3 made things &lt;em&gt;worse&lt;/em&gt; (more false&lt;br&gt;
positives, no real gain). The threshold is a fine-tuning dial, not a fix for bad training data.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 6 — The real jump: fine-tune with your own voice
&lt;/h2&gt;

&lt;p&gt;The synthetic-only model learned a TTS "Nova", not &lt;em&gt;my&lt;/em&gt; voice in &lt;em&gt;my&lt;/em&gt; room through &lt;em&gt;my&lt;/em&gt; microphone.&lt;br&gt;
The biggest, most reliable win was adding &lt;strong&gt;real recordings&lt;/strong&gt; of me saying the word and retraining.&lt;/p&gt;

&lt;p&gt;I wrote a small recorder that beeps and captures clips at 16 kHz mono, then varied distance, tone,&lt;br&gt;
speed and background noise across ~60 of them (and held ~10% back for evaluation). In the training&lt;br&gt;
notebook, those real positives get &lt;strong&gt;upsampled&lt;/strong&gt; and mixed in with the synthetic ones.&lt;/p&gt;

&lt;p&gt;The result, measured on the same evaluation set:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Recall went from 53% → 70% at threshold 0.35&lt;/strong&gt;, just by folding in real-voice positives.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not magic, but a real, measured step up — and exactly the kind of improvement that's invisible if&lt;br&gt;
you're only judging by ear.&lt;/p&gt;
&lt;h2&gt;
  
  
  Reproduce it yourself (the short version)
&lt;/h2&gt;

&lt;p&gt;Here's the minimal path with public tools, so you can train your own keyword. Swap &lt;code&gt;"nova"&lt;/code&gt; for&lt;br&gt;
yours throughout.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Install and synthesize positives.&lt;/strong&gt; Generate a few thousand clips, varying voice and speed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;piper-tts openwakeword
python &lt;span class="nt"&gt;-m&lt;/span&gt; piper.download_voices &lt;span class="nt"&gt;--download-dir&lt;/span&gt; voices &lt;span class="se"&gt;\&lt;/span&gt;
    es_ES-sharvard-medium es_MX-ald-medium es_MX-claude-high
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt;
&lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;positives&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;mkdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exist_ok&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;voices&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;es_ES-sharvard-medium&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;es_MX-ald-medium&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;es_MX-claude-high&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2000&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;choice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;voices&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;length&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uniform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1.3&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# speed variation
&lt;/span&gt;    &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;piper&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;voices/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.onnx&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--length-scale&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;length&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--output_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;positives/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.wav&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nova&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# (CLI flags vary slightly by Piper version — check `piper --help`.)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then &lt;strong&gt;listen to a handful&lt;/strong&gt; before going further.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Train.&lt;/strong&gt; Open openWakeWord's &lt;code&gt;automatic_model_training.ipynb&lt;/code&gt;&lt;br&gt;
(&lt;a href="https://github.com/dscripka/openWakeWord" rel="noopener noreferrer"&gt;in their repo&lt;/a&gt;) on Kaggle or Colab, select a &lt;strong&gt;T4&lt;/strong&gt;&lt;br&gt;
GPU, set &lt;code&gt;target_phrase = ["nova"]&lt;/code&gt;, point it at your &lt;code&gt;positives/&lt;/code&gt; folder, and &lt;em&gt;Run All&lt;/em&gt;. It pulls&lt;br&gt;
the negative/background datasets and augmentation for you and exports a single &lt;code&gt;nova.onnx&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Evaluate&lt;/strong&gt; against your own clip folders and sweep thresholds — this is the step people skip:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;soundfile&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;sf&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openwakeword.model&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Model&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wakeword_models&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nova.onnx&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;best_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reset&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sf&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                       &lt;span class="c1"&gt;# 16 kHz mono
&lt;/span&gt;    &lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;32767&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;astype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;int16&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;top&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1280&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;1280&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;    &lt;span class="c1"&gt;# 80 ms windows
&lt;/span&gt;        &lt;span class="n"&gt;top&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;top&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;audio&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1280&lt;/span&gt;&lt;span class="p"&gt;])[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nova&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;top&lt;/span&gt;

&lt;span class="n"&gt;pos&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;best_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eval/positives/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eval/positives&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
&lt;span class="n"&gt;neg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;best_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eval/negatives/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eval/negatives&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.35&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;recall&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;fp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;neg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thr=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: recall=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;recall&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;  false_positives=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;fp&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;neg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;4. Fine-tune with real voice (optional, high impact).&lt;/strong&gt; Record ~60 clips of yourself saying the&lt;br&gt;
word (16 kHz mono — a few lines with &lt;code&gt;sounddevice&lt;/code&gt;), hold back ~10% for evaluation, and add them to&lt;br&gt;
the training set with upsampling. This is what took me from 53% to 70%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Deploy.&lt;/strong&gt; Copy &lt;code&gt;nova.onnx&lt;/code&gt; to your device and feed the detector &lt;strong&gt;1280-sample windows&lt;/strong&gt; (see&lt;br&gt;
the buffering snippet above). Tune the threshold from your evaluation numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually moved the needle
&lt;/h2&gt;

&lt;p&gt;If you only remember four things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Match the TTS language to how you'll say the word.&lt;/strong&gt; This was worth more than any
hyperparameter.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Listen to your synthetic positives before training.&lt;/strong&gt; One minute saves hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure recall and false positives on a varied clip set.&lt;/strong&gt; Your ears lie; a threshold sweep
doesn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fine-tune with your own real voice.&lt;/strong&gt; Synthetic gets you started; real recordings get you
reliable.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And one bonus, because it bit me hardest: if a model scores zero on everything, check the window&lt;br&gt;
size you're feeding it before you retrain anything.&lt;/p&gt;

&lt;p&gt;Have you trained a custom wake word? I'd love to hear what moved recall for you — especially for&lt;br&gt;
non-English keywords.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Want more context, or to see how we did it?&lt;/strong&gt;&lt;br&gt;
Nova isn't fully public yet, but you can get &lt;strong&gt;early access&lt;/strong&gt; to the repository — all the&lt;br&gt;
documentation and the complete source code — at &lt;strong&gt;&lt;a href="https://gitlab.com/gabrielhruiz1/nova" rel="noopener noreferrer"&gt;https://gitlab.com/gabrielhruiz1/nova&lt;/a&gt;&lt;/strong&gt;.&lt;br&gt;
Leave us a message and we'll try to grant you access ASAP, until we publish everything&lt;br&gt;
officially (we're still working on a few parts).&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>tutorial</category>
      <category>raspberrypi</category>
    </item>
    <item>
      <title>Introducing Nova: an open-source voice assistant for the Raspberry Pi 4</title>
      <dc:creator>Gabriel Hidalgo</dc:creator>
      <pubDate>Thu, 08 Oct 2026 11:02:45 +0000</pubDate>
      <link>https://dev.to/gabrielhruiz/introducing-nova-an-open-source-voice-assistant-for-the-raspberry-pi-4-1d15</link>
      <guid>https://dev.to/gabrielhruiz/introducing-nova-an-open-source-voice-assistant-for-the-raspberry-pi-4-1d15</guid>
      <description>&lt;h1&gt;
  
  
  Introducing Nova
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Nova&lt;/strong&gt; is an open-source voice assistant that runs entirely on a &lt;strong&gt;Raspberry Pi 4&lt;/strong&gt;. You wake it&lt;br&gt;
with a keyword, speak to it in plain language, and it answers out loud. It stays quiet and local&lt;br&gt;
until you actually talk to it — only then does anything leave the device.&lt;/p&gt;

&lt;p&gt;This post is a tour of what Nova is and the handful of decisions that shaped it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline
&lt;/h2&gt;

&lt;p&gt;Everything hangs off one simple loop, governed by a small state machine&lt;br&gt;
(&lt;code&gt;IDLE → LISTENING → THINKING → SPEAKING → IDLE&lt;/code&gt;):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Wake word.&lt;/strong&gt; In its resting state Nova only listens &lt;em&gt;locally&lt;/em&gt; for the word "Nova", using a
small on-device model. Nothing is sent anywhere.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Listening.&lt;/strong&gt; Once woken, it records until you stop speaking, detected locally with voice
activity detection (VAD).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thinking.&lt;/strong&gt; It transcribes your speech to text (STT) and sends that text to a language model
(LLM) to produce an answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speaking.&lt;/strong&gt; It turns the answer into speech and plays it back. Then it goes back to sleep.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each stage sits behind an interface, so any one engine can be swapped without touching the rest.&lt;/p&gt;

&lt;h2&gt;
  
  
  The choices that matter
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Local where it counts, cloud where it helps.&lt;/strong&gt; Wake word, VAD, speech-to-text and the voice all&lt;br&gt;
run on the Pi. The language model runs in the cloud by default, simply because a 2 GB Pi can't host&lt;br&gt;
a strong LLM well. The LLM lives behind a provider interface, so you can point Nova at a hosted&lt;br&gt;
Claude, at AWS Bedrock, or at a local Ollama server on your LAN. A fully local setup on a more&lt;br&gt;
powerful node is on the roadmap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Privacy by construction.&lt;/strong&gt; Because the wake word runs on-device, Nova isn't streaming your room&lt;br&gt;
to the internet. Audio is transcribed and sent only when you are inside an active turn — never at&lt;br&gt;
rest. A short beep tells you when it is actually listening.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spanish-first, but configurable.&lt;/strong&gt; Nova was built and tuned for Spanish voice interaction,&lt;br&gt;
including a custom wake-word model trained on Spanish pronunciations (the English generators simply&lt;br&gt;
didn't match). Language and voices are configuration, not code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Robustness is sacred.&lt;/strong&gt; Nova is meant to run unattended for weeks. That pushed a lot of&lt;br&gt;
unglamorous work: recovering a USB mic that gets unplugged, a systemd watchdog that can tell a hung&lt;br&gt;
turn from a slow one, clean startup, and per-turn metrics so you can see what the box is doing —&lt;br&gt;
right down to reading the Pi's thermal throttling state.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why open source
&lt;/h2&gt;

&lt;p&gt;The whole stack is free and open. The only running cost of the first version is the LLM usage from&lt;br&gt;
whichever provider you choose. Everything else — the wake-word trainer, the installer, the design&lt;br&gt;
docs — is in the repository, and the design documentation is treated as the source of truth and&lt;br&gt;
kept in step with the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it's going
&lt;/h2&gt;

&lt;p&gt;Nova already works end to end and installs as a service. Next up: better documentation and a&lt;br&gt;
presentation site, a more natural voice, a performance benchmark with regression thresholds, and a&lt;br&gt;
fully local node so the whole pipeline can run without the cloud.&lt;/p&gt;

&lt;p&gt;If any of that sounds fun, come along for the ride.&lt;/p&gt;




&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Want more context, or to see how we did it?&lt;/strong&gt;&lt;br&gt;
Nova isn't fully public yet, but you can get &lt;strong&gt;early access&lt;/strong&gt; to the repository — all the&lt;br&gt;
documentation and the complete source code — at &lt;strong&gt;&lt;a href="https://gitlab.com/gabrielhruiz1/nova" rel="noopener noreferrer"&gt;https://gitlab.com/gabrielhruiz1/nova&lt;/a&gt;&lt;/strong&gt;.&lt;br&gt;
Leave us a message and we'll try to grant you access ASAP, until we publish everything&lt;br&gt;
officially (we're still working on a few parts).&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>raspberrypi</category>
      <category>python</category>
      <category>opensource</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
