A lot of what a developer types in a day is not code. It is commit messages, pull request descriptions, review comments, issue replies, design notes and Slack-length explanations. That text is ordinary prose, and prose is what speech recognition handles well. Syntax is what it handles badly.
This post is a practical split: what to dictate, what to keep on the keyboard, and how to set up a phone as the microphone for a Windows PC. I build the app at the end of it.
Disclosure: I'm the developer of Bolo, the app mentioned below, and this post is published from my account. It was written and published by an AI agent working on my behalf.
Dictate prose, type syntax
Speech engines are trained on sentences. They are good at turning "the retry loop now backs off before the third attempt" into clean text. They are poor at if (!cfg?.retries) return;. Trying to speak brackets, camelCase and operators is slower than typing them, and you will spend the saved time fixing them.
A split that works:
| Dictate | Keep on the keyboard |
|---|---|
| Commit message body | Commit subject line, if your team uses a strict prefix format |
| PR description and "how to test" steps | Code, config, shell commands |
| Review comments explaining why | Identifiers, file paths, version numbers |
| Issue replies and bug reports | Anything you will paste into a terminal |
| Docstrings and README paragraphs | Regexes, URLs |
The rule of thumb: if a human will read it as a sentence, speak it. If a machine will parse it, type it.
Habits that make dictated text need less fixing
- Think the sentence first, then say it in one go. Pausing mid-sentence is where most stray words and broken punctuation come from.
-
Leave identifiers out and add them after. Say "the function that parses the config", then replace it with the real name. That is faster than spelling
parseCfgV2aloud. - Dictate into a scratch place for long text. For a long PR description, speak into a plain editor, read it once, then paste. Editing in a web form you might accidentally submit is riskier.
- Read before you send. Recognition errors are plausible-looking words, not typos, so spell-check will not catch them.
Option 1: Windows' built-in voice typing
Windows has voice typing built in: press Win + H with the cursor in a text field. It needs no install, and for English with a decent microphone it is a good first thing to try.
Its limits depend on your hardware and your language. A laptop's built-in microphone a metre away in a noisy room gives poor results, and support for some languages is weaker than for English.
Option 2: use your phone as the microphone
Your phone's microphone is close to your mouth and is built for speech. Bolo uses that: you speak into an Android phone and the text is typed at the cursor on your Windows PC, over your Wi-Fi. The cursor can be in a browser form, an email, a document or a code editor.
What you need:
- an Android phone and a Windows PC on the same Wi-Fi network;
- Bolo on the phone (Google Play);
- the Bolo PC companion on Windows (GitHub releases).
Install both, pair them, put the PC cursor where you want text, and speak into the phone.
You choose how speech is recognised:
- Offline, on the phone: the audio does not leave the device. This is useful for anything under NDA or when the connection is poor.
- Bolo Cloud: dictation in 28 languages, including English, Hindi, Bengali and Telugu. It comes as a limited free cloud trial.
- Your own OpenAI or Gemini API key: you pay your provider directly.
There is an optional AI clean-up step, using OpenAI models, which adds punctuation and capitals and removes "umm" and "ahh". For commit messages and review comments that is most of the editing you would otherwise do by hand.
The clipboard half
The other half of the friction is moving text between devices: a stack trace someone sent you on your phone, an address, a link from a chat. Bolo has two-way clipboard sync between the phone and the PC, so you copy on one and paste on the other without emailing yourself.
Limits to know
- The PC side is Windows only for now.
- The phone and the PC need to be on the same Wi-Fi network.
- Offline recognition trades some accuracy for privacy. Cloud recognition needs an internet connection.
- None of this makes speaking code practical. Keep syntax on the keyboard.
Try the split for a week
Even without any app, try the table above for a week with Win + H: dictate every commit body and PR description, type everything else. If your wrists or your typing speed are the bottleneck, it adds up quickly.
If you try the phone-as-mic route, I'd like to hear which editor or language gave you trouble.
Top comments (0)