DEV Community

Cover image for Sidekick part 3: voice that never leaves the machine
Irfan Wani
Irfan Wani

Posted on

Sidekick part 3: voice that never leaves the machine

Sidekick Part 3: voice that never leaves the machine

Cloud voice assistants have a dirty secret everyone knows and nobody says: your voice goes to a server. Every "hey assistant" is an upload. Sidekick's sk talk does the whole loop on your hardware — OS-native capture, local transcription, transcript dropped editable into the prompt. Part 3 of this series is how that pipeline works, and the unglamorous bugs inside it.

The pipeline in 30 seconds

Enter starts recording, Enter stops it. Underneath: the OS-native recorder (arecord/ALSA on Linux, sox/CoreAudio on macOS, ffmpeg fallback) writes 16kHz wav, then local faster-whisper transcribes it — int8 quantized, CPU-only, model downloaded once and cached:

if model_size not in _model_cache:
    _model_cache[model_size] = WhisperModel(model_size, device="cpu", compute_type="int8")
model = _model_cache[model_size]
segments, _ = model.transcribe(wav_path, beam_size=5)
text = " ".join(s.text.strip() for s in segments).strip()
if not text:
    raise RuntimeError(
        "heard only silence — speak louder/closer, or run `sk mic-test` to check levels"
    )
Enter fullscreen mode Exit fullscreen mode

No hard dependencies: arecord ships with the OS, sox/ffmpeg come via brew, and faster-whisper installs itself into the running environment on first use (uv pip install faster-whisper, with a "still not importable — restart and retry" path for the inevitable edge case). In the TUI, ctrl+g or the mic pill does the same thing. --stt-model base trades accuracy for speed when words vanish.

One technical decision: errors written for the speaker, not the log

Look at the failure messages: "recording is nearly empty — mic may be muted, run sk mic-test." "Heard only silence — speak louder/closer." Each one names the likely physical cause and the next command. That came from watching real sessions: transcription failures are almost never model failures, they're room failures — muted mics, wrong devices, whisper-quiet laptops. So sk mic-test records three seconds and returns a verdict — silent, quiet, good — with peak dB and an alsamixer hint when it's the OS mixer lying to you.

The privacy side is structural, not promised. Recordings are temp files deleted after every take — there's nowhere for them to accumulate because nothing in the pipeline has a server to send them to. The strongest privacy guarantee is an architecture with no upload path.

What broke

Two favorites from src/sk/voice.py. First: Python 3.13 removed audioop, so mic-level measurement needed a hand-rolled PCM stats function with struct — including a comment about endianness placement in format strings that reads like someone lost an afternoon to it.

Second, the spawn guard. faster-whisper's dependency tree can start a multiprocessing child, and on macOS recycled file descriptors made CPython reject the spawn with ValueError: bad value(s) in fds_to_keep. The fix monkeypatches spawnv_passfds to sanitize the fd list — dedupe, sort, drop closed ones — installed once and only around STT work. And because callers reduce failures to one-line messages, which hid the genuine crash site, full tracebacks now append to ~/.sidekick/voice-errors.log. The bug that teaches you your error reporting is broken is always two bugs.

Part 4 goes into the tool safety model: 19 tools, approval gates, hard refusals, and the audit ledger.


Built by the Sidekick community. Repo: https://github.com/Faisal-Fayaz/sidekick

Top comments (0)