I built YazSes: hold a key, speak, release, and your words are typed into whatever app has focus. It runs speech-to-text locally (faster-whisper on CPU), so nothing leaves your machine: no cloud, no account, no API key.
Install
Linux
curl -fsSL https://raw.githubusercontent.com/MSKazemi/yazses/main/install.sh | bash
Snap
sudo snap install yazses
sudo snap connect yazses:audio-record # microphone
sudo snap connect yazses:raw-input # hold-to-talk key
macOS (Apple Silicon)
brew tap MSKazemi/yazses && brew trust MSKazemi/yazses && brew install --cask yazses
Windows
winget install MSKazemi.YazSes
Any OS with Python 3.11+
pipx install yazses
Then:
yazses quickstart # 3 steps tailored to your machine, read-only
yazses doctor # verify microphone, hotkey and text injection
Why offline?
- Your voice never leaves the machine.
- No subscription, no API key, works without internet.
- It also transcribes recordings and captures meetings with speaker labels, locally.
Links
- GitHub (stars welcome): https://github.com/MSKazemi/yazses
- Install guides: Linux · macOS · Windows
- Docs: https://mskazemi.com/yazses
Which OS are you on? Tell me what breaks and I'll fix it.
Want to contribute? Here is where you can help in 15 minutes
You do not need to be a speech-recognition expert. Real, open tasks right now:
- 🌍 Native speaker of Persian, Urdu, Thai, Vietnamese, Telugu, Ukrainian, Traditional Chinese...? Several translations of the docs have never been read by a native speaker. Fixing a sentence is a valid first PR.
- 🧪 Pick a
good first issue: https://github.com/MSKazemi/yazses/labels/good%20first%20issue - 🐧 Test it on your distro / desktop (Wayland, X11, Windows, macOS) and open an issue with the output of
yazses doctor. Bug reports from real machines are gold. - 📖 Read the guide: https://github.com/MSKazemi/yazses/blob/main/.github/CONTRIBUTING.md
The test suite runs offline (no microphone, no model download), so you can run it in a minute:
uv sync
uv run python -m pytest tests/ -v
Try it, then tell me
- ⭐ Star the repo so others can find it: https://github.com/MSKazemi/yazses
- ⬇️ Install it with the one-liner above and run
yazses doctor. - 💬 Comment below with your OS and what happened. I reply to every comment.
Top comments (1)
For "typed into whatever app has focus," I'd add a delayed-transcription focus-change fixture: record in editor A, release the key, then move focus to app B before inference finishes. Repeat with A closed and with the original field no longer editable.
The key assertion is that text doesn't silently land in B just because B has focus when recognition completes. A recoverable preview or explicit insert action would make an uncertain target visible without losing the transcript. Does YazSes bind insertion to the recording target or resolve focus at completion? I haven't installed it; this is a suggested text-injection test from the interaction described here, not a report of a bug.