I built an open-source voice input for macOS that never uploads your voice
Most voice input tools upload every recording to somebody's server. I wanted something simpler: hold a hotkey, talk, release — and the text appears at my cursor, with my voice never leaving the Mac.
So I built Cadenza (随言). MIT open source, macOS 14+, Apple Silicon and Intel.
- Repo: https://github.com/DragonKingIO/Cadenza-voice
- Site & docs: https://dragonkingio.github.io/cadenza-site/
Why local-first
This category has a trust problem. We've watched popular dictation apps regress in accuracy after updates, get caught in screenshot scandals, and jack up lifetime prices from $249 to $849. Every one of them routes your voice through their servers.
Cadenza's answer is architectural: recognition happens on your machine, period. There's no server to breach, no recording to subpoena, no subscription to hike.
How it works
1. Hold to talk. Press and hold a shortcut (configurable; tap-to-toggle also works). A small recording bar shows what's happening — recording, recognizing, done.
2. Recognize on-device. Five models to choose from — SenseVoice (my recommendation), FireRedASR2, Paraformer, Parakeet, Qwen3-ASR — running locally via sherpa-onnx. Models download inside the app with checksum verification, and inference needs no network connection at all. The list shows each model's speed, memory footprint, and punctuation support so you can pick your tradeoff.
3. Text lands at your cursor. Release the key and the transcription is typed wherever your cursor is. Filler words ("uh", "um") and repeated characters are cleaned up by default; an optional AI polish step fixes punctuation and obvious errors — text only, and it can run against a local Ollama instance.
Privacy is a design, not a promise
- No account, no telemetry, no analytics.
- Recordings and transcripts are never written to disk and never logged.
- API keys live in the macOS Keychain, not in a config file.
- There's even a "never go online" kill switch.
The code is MIT open source, so don't take my word for it — you can audit exactly what leaves your computer.
Cloud, strictly on your terms
If you do want a cloud provider, you bring your own API key, and each provider receives audio only after you authorize it individually. Supported: OpenAI, Groq, Google, Azure, AssemblyAI, ElevenLabs, plus any OpenAI-compatible endpoint. These also power the optional voice-translation hotkey (separate shortcut, types the translation directly) and the AI polish step.
One honest disclosure: I haven't tested the cloud providers against real paid accounts yet. If you try one, I'd genuinely like to hear how it went.
Made for real vocabularies
There's a vocabulary manager for your proper nouns, project names, and jargon — plus shared packs you can import. This is the unglamorous feature that actually determines whether dictation is usable for work.
Also in testing: screenshot OCR (select a region, read text with on-device Apple Vision or PP-OCR).
Status: early preview, daily-driven
We use it every day, but it hasn't been tested in many environments yet. Two practical notes:
- It's not notarized by Apple yet, so first launch needs System Settings → Privacy & Security → "Open Anyway", plus microphone, accessibility, and input monitoring permissions.
- Requirements: macOS 14+, Apple Silicon or Intel.
What would help most right now: recognition feedback across accents, microphones, and real rooms; real-world cloud provider results; UI translations and docs in other languages.
If it's useful, a star is appreciated — and if not, that's completely fine too. Issues and PRs welcome.







Top comments (2)
Nice to see the local-only path done properly. One thing that will bite early-preview users of a non-notarized app that needs Accessibility and Input Monitoring:
tccutil reset Accessibility <your.bundle.id>(same forListenEvent= Input Monitoring), then grant again.AXIsProcessTrusted()and show a clear "permission missing" state in the recording bar instead of letting a release produce nothing. Most "it stopped working after the update" issues for apps like this are exactly that.If text insertion goes through the clipboard plus Cmd+V, restoring the previous clipboard contents afterwards is worth it too, otherwise every dictation overwrites whatever the user had copied.
We ran into the same permission trap at Auten, where an AI agent drives the Mac through the accessibility tree.
tr.ee/dev-to