DEV Community

Mohsen Seyedkazemi Ardebili
Mohsen Seyedkazemi Ardebili

Posted on

I got tired of typing, so I built a free, offline voice-typing tool (open source, contributors wanted)

I built YazSes: hold a key, speak, release, and your words are typed into whatever app has focus. It runs speech-to-text locally (faster-whisper on CPU), so nothing leaves your machine: no cloud, no account, no API key.

Install

Linux

curl -fsSL https://raw.githubusercontent.com/MSKazemi/yazses/main/install.sh | bash
Enter fullscreen mode Exit fullscreen mode

Snap

sudo snap install yazses
sudo snap connect yazses:audio-record   # microphone
sudo snap connect yazses:raw-input      # hold-to-talk key
Enter fullscreen mode Exit fullscreen mode

macOS (Apple Silicon)

brew tap MSKazemi/yazses && brew trust MSKazemi/yazses && brew install --cask yazses
Enter fullscreen mode Exit fullscreen mode

Windows

winget install MSKazemi.YazSes
Enter fullscreen mode Exit fullscreen mode

Any OS with Python 3.11+

pipx install yazses
Enter fullscreen mode Exit fullscreen mode

Then:

yazses quickstart   # 3 steps tailored to your machine, read-only
yazses doctor       # verify microphone, hotkey and text injection
Enter fullscreen mode Exit fullscreen mode

Why offline?

  • Your voice never leaves the machine.
  • No subscription, no API key, works without internet.
  • It also transcribes recordings and captures meetings with speaker labels, locally.

Links

Which OS are you on? Tell me what breaks and I'll fix it.

Want to contribute? Here is where you can help in 15 minutes

You do not need to be a speech-recognition expert. Real, open tasks right now:

The test suite runs offline (no microphone, no model download), so you can run it in a minute:

uv sync
uv run python -m pytest tests/ -v
Enter fullscreen mode Exit fullscreen mode

Try it, then tell me

  1. ⭐ Star the repo so others can find it: https://github.com/MSKazemi/yazses
  2. ⬇️ Install it with the one-liner above and run yazses doctor.
  3. 💬 Comment below with your OS and what happened. I reply to every comment.

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

For "typed into whatever app has focus," I'd add a delayed-transcription focus-change fixture: record in editor A, release the key, then move focus to app B before inference finishes. Repeat with A closed and with the original field no longer editable.

The key assertion is that text doesn't silently land in B just because B has focus when recognition completes. A recoverable preview or explicit insert action would make an uncertain target visible without losing the transcript. Does YazSes bind insertion to the recording target or resolve focus at completion? I haven't installed it; this is a suggested text-injection test from the interaction described here, not a report of a bug.