DEV Community

Alexandre Machado
Alexandre Machado

Posted on

I Gave Claude Code a Voice on My Windows Laptop. The Slowest Part Is the Voice.

At 18:59 on a Saturday I asked my laptop, in Portuguese, for a short summary of the most important tech news of the week. A voice I had cloned said "Um instante" ("one moment"). Claude Code ran a web search, and about 17 seconds after I stopped talking her answer started coming back as speech. Her pick was Google putting the first prototype satellite of Project Suncatcher in orbit.

Nothing in that loop typed a word. Speech recognition, the voice and the audio pipeline all run on the laptop. The only thing that leaves the machine is the text Claude Code sends to its own API.

The demo is in Portuguese, my first language. Turn on English captions.

AI disclosure: most of the code in this project was written by coding agents (Claude Code, OpenAI Codex and Google Antigravity) under my direction. I drafted this article with Claude from the repo's commit history, design notes and logs, then edited it myself. Every number below comes from those logs.

Where Débora came from

Débora started as a dictation tool. I forked npu-whisper, a Windows app that runs Whisper on Intel's NPU, and turned it into something I use every day: press Ctrl+Space, talk, and the text lands at the cursor. I added NVIDIA support through faster-whisper, a device priority list with fallback (RTX, then NPU, then iGPU, then CPU) and a uv tool install setup so it installs as one command instead of an afternoon of virtualenvs.

The laptop matters for everything that follows: a Core Ultra 9 185H (with an NPU and an Arc iGPU) and an RTX 4070 Laptop GPU with 8 GB of VRAM.

When I wrote about it on LinkedIn, I said the next step was connecting this voice layer to the AI agents I already use in the terminal. About 30 hours and 55 commits later, that step is done. These are the decisions that shaped it, and the place where it is still slow.

Decision 1: a local LLM was good enough to chat, and I still didn't use it

The first version of voice chat ran a local model as the brain. Before replacing it, I benchmarked three on the Arc iGPU with OpenVINO GenAI, int4, same Portuguese conversations, three samples each:

Model First token Tokens/s Followed the long persona prompt
Qwen3-8B 0.3 s ~11 Badly: misread my name, wrote dates as digits, used emoji
Qwen3.5-9B 2.1 s ~6 Yes, but about 4x slower to the first token
Gemma 4 E4B 0.2 to 0.4 s ~15 Yes, every rule, and the only one that used feminine Portuguese forms on its own

Gemma 4 E4B won clearly. It answered "Débora, que dia é hoje?" with the full date spelled out for the speech engine, and asked me to repeat myself when the transcription came out garbled. With Qwen3-8B, every rule I added to the prompt made the answers worse.

I went with Claude Code anyway. A local 4B model can hold a pleasant conversation, but it can't open my repositories, check a container, read the app's own logs or search the web. Claude Code already has the tools, the project context and my instructions. So the split became: Débora owns the ears and the mouth (Whisper and the TTS), and a harness I didn't write owns the thinking. The local model is still in the code as a backend for when I want everything offline.

The obvious cost is latency. A cold claude -p call took 6.1 s for a short question. I measured the other CLIs the same way:

CLI Version Cold call
claude -p 2.1.295 6.1 s
codex exec 0.161.0 6.1 s
copilot -s -p 1.0.93 7.9 s
agy -p 1.3.2 9.3 s

Six seconds of silence after every sentence would kill a voice conversation. That led to the next decision.

Decision 2: keep one Claude Code process alive

Claude Code can run as a long-lived process that reads and writes JSON lines over stdio. This is roughly how Débora starts it:

claude -p --input-format stream-json --output-format stream-json --verbose \
  --include-partial-messages \
  --session-id <uuid> \
  --permission-mode acceptEdits \
  --permission-prompts host --permission-prompt-tool stdio \
  --append-system-prompt-file <voice-rules.md>
Enter fullscreen mode Exit fullscreen mode

Each flag solves one voice problem:

  • --include-partial-messages streams text as it is generated, so the TTS can start on the first phrase instead of waiting for the full answer.
  • --session-id and --resume keep the same conversation across app restarts. I can also open that session in a normal terminal with claude --resume <id> and see everything she did.
  • --permission-prompt-tool stdio turns every tool permission into a control_request with subtype: "can_use_tool" on stdout. Débora answers on stdin.
  • --append-system-prompt-file adds the voice-channel rules (short spoken answers, no markdown, numbers and dates written out, the date and language of the session) on top of whatever CLAUDE.md the project already has.

The permission answer is small:

response = ({"behavior": "allow", "updatedInput": request["input"]} if allow else
            {"behavior": "deny", "message": "Permissão negada pela Débora."})
return {"type": "control_response", "response": {
    "subtype": "success", "request_id": request_id, "response": response}}
Enter fullscreen mode Exit fullscreen mode

The default is to deny, with an allowlist for read-only diagnostics. For the demo I had opened it up, which is why the log shows Claude pede permissão: WebSearch followed by permission allowed.

Before writing any of this into the app, I proved it with a throwaway script. Two turns in the same process kept a test word ("jabuticaba") in context. The first text delta took 3.8 s cold and 1.8 s warm. An interrupt sent as a control_request with subtype: "interrupt" ended the turn with terminal_reason: "aborted_streaming", and the process stayed alive. That last part matters, because people talk over a voice assistant all the time.

Decision 3: send the raw transcription, mistakes included

Whisper hears "dictation engine" in a Brazilian accent as "dictêixon engine". My first idea was to have a local LLM clean the transcription before it reached Claude. I tested it, and the cleanup made things worse. It fixed the obvious errors, then "corrected" the project-specific ones into something wrong ("dictêixon engine" became "Decision Engine"), and it added about a second.

Claude Code understood all four raw test requests, because it had the repository open and could see dictation_engine.py sitting right there. So Débora sends the raw text, and the system prompt says it came from speech recognition and may contain errors.

The fix for recognition errors went to the other end of the pipeline:

  • Claude keeps a small voice memory at ~/.debora/harness/voice_memory.md, one line per correction: - "what was heard" → correct term (optional context). Only sessions started by Débora get this file in their prompt.
  • The terms from that file, plus the project names found in the working folder (folder name, pyproject.toml, package.json), are passed to Whisper as hotwords, up to 40 terms.

Over time the recognizer picks up my project vocabulary, and fewer mistakes reach Claude in the first place.

Decision 4: never lose what the user said

The bug that annoyed me most in testing: I paused mid-sentence, Débora started answering, and the second half of what I said disappeared. The fix touched several parts of the pipeline:

  • Whisper re-transcribes the growing segment about once a second while I'm still talking (in the demo, 11.9 s of speech transcribed in 0.2 s on the RTX, a real-time factor of 0.02). If the draft looks unfinished, like a trailing "and..." or an ellipsis, the VAD waits 2 s of silence instead of 0.8 s before cutting.
  • Capture and transcription keep running while Claude thinks and while Débora speaks. Nothing gets muted.
  • Replies go through a serial queue. Barge-in can interrupt the active turn, drain the rest of Claude's response and send the next utterance in order.
  • An echo filter compares what the mic heard against the sentences Débora was actually playing (60% token overlap means echo), so on laptop speakers she doesn't answer herself.

The numbers, from the demo log

Step Turn 1: "Oi Débora, tudo bem?" Turn 2: tech news
Final transcription after I stopped ~0.2 s ~0.2 s
Claude first token 1.4 s 17.2 s (ToolSearch + WebSearch)
First audio 31.8 s ~4.8 s ("Um instante")
Whole reply spoken 49.2 s still talking when the log excerpt ends, 52 s in

Claude was fast on the first turn and audio was slow. The first sentence was queued at 2.3 s, but synthesis didn't start until 28 s. VRAM went from 1.3 GB to 5.0 GB in that window, so my reading is that the TTS server was still loading its model. The fix is to start it earlier and warm it up before the first turn.

The second turn shows the real problem. Débora's voice is a clone made with Chatterbox Multilingual from a short reference recording, and on this GPU it synthesizes slower than real time. The chunks in that reply took between 1.1x and 1.4x their own audio length to generate (for example, 5.9 s of work for 5.1 s of speech). Every sentence waits for the previous one, so the gap grows over a long answer.

Claude Code was not the bottleneck. With the process warm, its first token came in 1.4 s on a simple turn, and the web search turn spent its time on tools. What slows the conversation down is the voice: a voice-cloning TTS model sharing 8 GB of VRAM with Whisper on a laptop.

Who actually wrote the code

Most of the 55 commits were written by agents. Claude Code coordinated: it broke the work down, wrote most of the harness and dispatched tasks. Codex picked up work in parallel through codex exec, mostly because I had spare quota there, including the code and security reviews before merging. Antigravity got research questions and UI ideas, like how to show voice mode in both the tray icon and the overlay. Its first pass left some mess behind that Claude later cleaned up.

My part was the one the agents couldn't do. I talked to her while testing, found bugs by ear ("your text gets cut off next to the mascot", "half of my sentence vanished when you interrupted me"), read the logs with them, and decided what stayed. Some of those decisions were about what not to build: no pre-cleanup LLM, no terminal UI for now, and no new Windows-only dependencies.

Limitations

  • Débora runs her own headless Claude Code session. She can't drive a terminal you already have open.
  • Claude answers in markdown. Only the speakable part goes to the TTS, and code and lists stay in the overlay and the log.
  • It is Windows-only today, and the voice chat needs a GPU to be pleasant.
  • The Whisper language setting matters. If it's set to English, Portuguese speech gets translated before Claude sees it.

What's next

The next piece is automatic hardware distribution. At startup Débora will probe the NPU, the iGPU and the RTX, then decide where Whisper, the TTS and a local model should each run, with a fallback for each component instead of one global device. The goal is to take the voice off the critical path. One option on this laptop is Whisper on the NPU, which would leave the whole RTX to the TTS.

The code is MIT licensed: github.com/alexandre-machado/debora-whisper.

One question, because I haven't solved it: has anyone gotten a local, voice-cloned TTS running faster than real time on 8 GB of VRAM? Which model, and what did it cost you in quality?

Top comments (2)

Collapse
 
reidmarlow profile image
Reid Marlow •

The autoregressive architectures used in Chatterbox and similar zero-shot cloners struggle to break 1.0 RTF on laptop GPUs because every acoustic token is an iterative generation step. On an 8 GB card, F5-TTS or E2-TTS using non-autoregressive flow matching drops the RTF down to around 0.3 to 0.4 once warmed up, but you still need Whisper completely off the discrete GPU. Moving Whisper to your Core Ultra NPU via OpenVINO is the right call there, because sharing VRAM between Whisper and a 5 GB TTS context causes memory bandwidth throttling even when you do not hit an outright OOM. If exact zero-shot timbre match can be relaxed to voice style blending, Kokoro runs at roughly 0.05 RTF on negligible VRAM, but for true reference-wav cloning on a laptop 4070, non-autoregressive flow matching is about the only way to avoid that queue backlog.

Collapse
 
suppdevbot profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to

Some comments have been hidden by the post's author - find out more