DEV Community

Krish Verma
Krish Verma

Posted on

How I built hands-free voice mode for my AI assistant: ffmpeg VAD, barge-in, zero npm packages

How I built hands-free voice mode for my AI assistant: ffmpeg VAD, barge-in, zero npm packages

A voice interface sounds like it needs a speech SDK, a streaming framework, and a small pile of native dependencies. Mine is one 1,100-line JavaScript module, zero npm dependencies, and it runs on binaries you probably already have: ffmpeg and ffplay.

The module is src/channels/voice.mjs in the Ankita repo. Here's how each piece works.

One ffmpeg process does recording and voice detection

The trick I like most: I never wrote a VAD (voice activity detection) library. ffmpeg ships a silencedetect audio filter, so a single ffmpeg process records the mic and streams speech boundaries on its stderr:

silencedetect=noise=-35dB:d=1.2
Enter fullscreen mode Exit fullscreen mode

The mic is captured as 16kHz mono PCM (-ac 1 -ar 16000 -c:a pcm_s16le), which is exactly what Whisper wants, and meanwhile stderr carries lines like silence_start: 3.4 / silence_end: 5.1 | silence_duration: 1.7. A small stateful parser (createSilenceParser) turns those into events, and createTurnDetector maps them onto a tiny state machine: sound-after-silence means "speech", silence-after-sound means "end" — the end of one spoken utterance.

The same turn detector is shared by hands-free capture and barge-in, so there's one source of truth for "is the human talking now?"

Speech-to-text and text-to-speech, no SDKs

  • STT goes to Groq's Whisper API (whisper-large-v3-turbo). One fetch call with multipart audio. It needs a GROQ_API_KEY, and that's the only key in the whole voice path.
  • TTS defaults to Microsoft Edge's free endpoint — and this is the part I find most satisfying. Instead of adding a WebSocket library, I did the handshake over Node's built-in node:tls, generating the Sec-MS-GEC token the endpoint expects and speaking the raw WebSocket protocol by hand. Thirty-or-so lines, no ws package.

If a Groq key is present you can flip to Orpheus TTS instead (canopylabs/orpheus-v1-english, with nine named voices, default tara). The "auto" provider setting picks Groq when a key exists and Edge otherwise. Long replies are split into sentences of at most 190 characters, synthesized chunk by chunk, then concatenated as WAV — Edge's endpoint is happier with short inputs.

Barge-in: interrupting the assistant mid-sentence

Real conversation means being able to interrupt. While a reply is playing through ffplay, the mic stays open and the same turn detector keeps feeding events. If speech starts more than ~400ms after playback began, that's a barge: kill the player with SIGKILL, bump a playbackGeneration counter so any audio that's still being written to disk can't play over the new utterance (a real race I hit — the file lands a moment after you kill the player), then transcribe what you said and feed it straight back into the loop as the next turn.

The 400ms grace window is deliberate: the first fraction of a second of "speech" after playback starts is often the assistant hearing its own tail through the mic.

The Bluetooth gotcha that nearly shipped

This one cost me an evening. Bluetooth headsets drop their high-quality A2DP audio endpoint the moment the mic opens — so if playback was routed through A2DP, barge-in replies came out silent. The fix lives in src/channels/audio-device.mjs: when barge-in is on and the mic is a Bluetooth device, playback is re-routed to the headset's Hands-Free endpoint before voice mode starts, then restored on exit. Not glamorous, but without it the whole feature is broken on the most common hardware.

It doesn't read you the whole reply

Long tool outputs don't work as speech. By default the assistant summarizes what it did into a short spoken note — at most 8 items, 6 sentences, 1200 characters — then tells you where the rest lives:

"{address}, that's {spoken} of {total}. The rest is on your screen, {address}."

(The default address is "sir" — configurable, and yes, the template tidies up its own punctuation if you leave it empty.) You can ask for a full read with VOICE_FULL_READ, but I kept it off by default because full reads are how you end up listening to a 2,000-token npm error trace.

What I'd do differently

The VAD is a blunt instrument: a fixed −35dB threshold can't tell keyboard clicks from speech, so the noise level, silence window, barge threshold and max utterance length are all tunable (VOICE_NOISE_DB, VOICE_SILENCE_MS, VOICE_BARGE_DB, VOICE_MAX_UTTERANCE_MS) instead of something clever. A proper ML-based VAD would be the real fix — but it would also be a native dependency, and the whole point was staying dependency-free. Sometimes "works on any machine with ffmpeg installed" beats "smart".

It's covered by 30 unit tests in test/channels/voice.test.mjs — the parsers and the turn detector are pure functions, which made them easy to test without ever touching a microphone.

The takeaway: a convincing voice mode doesn't require a speech platform. It requires one ffmpeg process, two API calls, a 4-state turn machine, and a healthy respect for Bluetooth.

If you build something with this, or you spot the edge case I missed — I'd genuinely like to hear about it in the comments.

Top comments (0)