The dictation tools I tried on my Mac fell into two camps. The built-in one punctuates poorly. The good ones are paid, and they send your voice to a remote service.
I wanted something smaller: press a shortcut, talk, and have the text appear in whatever app I'm in. No account, no subscription, no API key, nothing leaving the machine.
So I built Murmure. Version 1.1.0 is the first public release: github.com/croustibat/murmure.
What it does
Murmure is a menu bar app. You press Cmd+Shift+E, a small floating pill appears while it listens, you press the shortcut again, and the text is transcribed and pasted into the app that was active when you started. Mail, Slack, a terminal, an editor, a search field: it doesn't care.
Transcription runs locally with whisper.cpp and the large-v3-turbo model. It takes about two seconds between the end of a sentence and the pasted text.
There are two modes and no setting to switch between them:
- Short press: toggle. Press to start, press again to transcribe.
- Hold: push-to-talk. Keep the shortcut down while you speak, release to transcribe. The threshold between the two is 600 ms.
Escape cancels a dictation while it's listening. The recording is thrown away and nothing is pasted. Outside of a dictation, Escape keeps its normal role everywhere.
The menu also keeps your last 10 dictations. One click copies one back to the clipboard, which is handy when a paste fails.
The interesting bug: a script has no identity
At its core, Murmure is a shell script: it records with ffmpeg, sends the WAV to whisper-cli, and pastes the result. The obvious setup is to bind that script to a shortcut and call it a day.
Do that, and you mostly get subtitle credits.
The reason is macOS privacy permissions (TCC). A bare script has no identity that macOS can attach a microphone permission to. So macOS never asks for permission, and it doesn't return an error either: it hands the process a silent audio stream. Whisper, hearing nothing, does what it was trained to do on silent video and invents closing credits.
The fix is the reason Murmure is an .app bundle and not just a script. The bundle gets its own identity, macOS shows the microphone prompt, and the audio is real.
The pipeline today looks like this:
Cmd+Shift+E → Murmure.app (menu bar) → murmure.sh
├─ ffmpeg (avfoundation) ──→ WAV 16 kHz mono
├─ overlay (AppKit) ───────→ floating pill
├─ whisper-cli ────────────→ text
├─ corriger.pl ────────────→ vocabulary fixes
└─ pbcopy + Cmd+V ─────────→ active app
The same identity question comes back with the Accessibility permission, which Murmure needs to send Cmd+V. macOS ties that permission to the app's signature. If you have an Apple developer certificate in your keychain, install.sh signs the app with it and the permission survives updates. Without one, the app is signed ad hoc and recognized by its binary hash, so every rebuild means granting Accessibility again. The installer tells you when that happens, and it only rebuilds the app when something actually changed.
Murmure also checks both permissions itself. If one is missing, the menu bar icon shows an exclamation mark and the first menu entry opens the right System Settings panel.
Fixing technical words
I dictate in French, and Whisper likes to translate developer jargon into French words that sound close. "Commit" comes back as "commis", "worktree" as "worktrade".
Murmure has two files for that, both in ~/.local/share/murmure/:
-
vocabulaire.txtis a short text given to Whisper as a prompt before transcription. It's not a list to fill up: Whisper re-reads it before every 30-second chunk and it shares 224 tokens with the end of the previous chunk. Past roughly 150 tokens, long dictations start losing the thread between chunks. -
corrections.txtis applied after transcription, one rule per line, case-insensitive, on whole words:
commis|commit
redit|Redis
worktrade|worktree
In practice the corrections file does most of the work. The vocabulary prompt is best kept for proper nouns.
Speed, and what didn't help
I spent some time trying to make transcription faster. On a short dictation the encoder still processes a full 30-second window, and that's most of the time. Shrinking that window (-ac) made Whisper repeat itself. Greedy decoding, --no-fallback, thread counts and VAD gained less than 10% or hurt the transcription. So Murmure keeps whisper-cli defaults.
If you want to measure on your own recordings, the repo includes a small benchmark script. Put a reference .txt next to each .wav and it also computes the error rate:
scripts/bench.sh -n 5 dictee1.wav dictee2.wav
Network
Transcription never leaves the Mac. The only network request is an update check, at most once a day, that reads the latest release from the public GitHub API. No text, audio or identifier is sent. MURMURE_CHECK_UPDATES=0 in the config file turns it off.
Try it
You need an Apple Silicon Mac on macOS 14 or later, Homebrew, and the Xcode command line tools.
git clone https://github.com/croustibat/murmure.git
cd murmure
./install.sh
The installer sets up ffmpeg and whisper-cpp, downloads the model once (about 550 MB, checked by SHA-256), builds the app and launches it. To update: git pull && ./install.sh.
A note for non-French speakers: the docs and the menus are in French for now, and the default transcription language is French. Set MURMURE_LANG=en (or auto) in ~/.local/share/murmure/config and it applies to the next dictation.
It's MIT licensed. Issues and PRs are welcome, and I'd like to hear how it behaves with other languages and other microphones: github.com/croustibat/murmure.
Baptiste — freelance web dev for 20 years, building in Laravel/PHP and shipping small open-source tools.
Top comments (0)