I built VoxScribe, a free, open source voice dictation app for Windows. You hold a hotkey anywhere, talk, release, and the text is typed into whatever app has focus. Speech recognition runs locally with faster-whisper, so the audio never leaves the machine.
The Whisper part was the easy part. Almost every real bug came from audio hardware and drivers. Here are the ones that cost me the most time.
The first was a Bluetooth headset that recorded near silence. On my machine the sounddevice library picked a default input device that landed on an MME driver, and for my headset that produced audio so quiet that Whisper transcribed almost nothing. The same physical microphone shows up several times, once for each Windows audio host API (MME, WASAPI and DirectSound), and they are not interchangeable. The fix was to ask for the WASAPI default input device explicitly instead of trusting the generic default. Same headset, usable level right away.
The second was a crash when forcing 16 kHz. Whisper wants 16 kHz audio, so my first version opened the microphone at 16 kHz. That worked on my headset, then failed with an invalid sample rate error on a laptop's internal mic running at 48 kHz. Some drivers just refuse a rate they do not support natively. The fix was to record at the device's own native rate and resample to 16 kHz in software afterward. Plain linear interpolation is not audiophile quality, but for speech going into Whisper it was good enough, and it has worked on every device I have tried so far.
The third was dropping voice activity detection. My original design was fully automatic: run a voice activity detector (Silero VAD) on the live stream, start a segment when it hears speech and stop when it hears silence. On quiet or variable input, like a compressed Bluetooth headset, the detector's confidence was inconsistent. Worse, false triggers did not just miss words, they sent background noise to Whisper, and Whisper does not fail quietly on noise. It makes up plausible sentences. I switched to hold-to-talk, so the user decides when speech starts and stops and there is nothing to guess. I kept the VAD code in the repo but it is no longer used in the live path. Letting faster-whisper's built-in VAD trim silence inside a finished clip turned out to be enough.
There is also a known weakness in any design that opens the mic when you press a key: the stream takes a moment to start, especially on Bluetooth, so the first word can get clipped. I added an opt-in mode that keeps the stream open and holds the last half second in a small in-memory buffer, so a recording starts with audio that is already there. It is off by default, because an always-open mic means Windows shows the mic as in use all the time. Nothing is saved or transcribed unless you press the hotkey.
If I had to give advice to anyone building something like this: test on at least three different microphones before you trust anything, never assume a sample rate, record at the native rate and resample after, and expect the AI part to be the smallest source of bugs.
The project is MIT licensed if you want to read the capture code or try it: https://github.com/ahmedhmam1994/voxscribe-ai-voice-dictation. It is Windows only for now, since that is the machine I can test on. I used AI coding assistance while building it, and I tested every fix on real hardware myself. Issues and feedback are welcome, especially from people with unusual microphones.
Top comments (0)