Sidekick Part 3: voice that never leaves the machine
Cloud voice assistants have a dirty secret everyone knows and nobody says: your voice goes to a server. Every "hey assistant" is an upload. Sidekick's sk talk does the whole loop on your hardware — OS-native capture, local transcription, transcript dropped editable into the prompt. Part 3 of this series is how that pipeline works, and the unglamorous bugs inside it.
The pipeline in 30 seconds
Enter starts recording, Enter stops it. Underneath: the OS-native recorder (arecord/ALSA on Linux, sox/CoreAudio on macOS, ffmpeg fallback) writes 16kHz wav, then local faster-whisper transcribes it — int8 quantized, CPU-only, model downloaded once and cached:
if model_size not in _model_cache:
_model_cache[model_size] = WhisperModel(model_size, device="cpu", compute_type="int8")
model = _model_cache[model_size]
segments, _ = model.transcribe(wav_path, beam_size=5)
text = " ".join(s.text.strip() for s in segments).strip()
if not text:
raise RuntimeError(
"heard only silence — speak louder/closer, or run `sk mic-test` to check levels"
)
No hard dependencies: arecord ships with the OS, sox/ffmpeg come via brew, and faster-whisper installs itself into the running environment on first use (uv pip install faster-whisper, with a "still not importable — restart and retry" path for the inevitable edge case). In the TUI, ctrl+g or the mic pill does the same thing. --stt-model base trades accuracy for speed when words vanish.
One technical decision: errors written for the speaker, not the log
Look at the failure messages: "recording is nearly empty — mic may be muted, run sk mic-test." "Heard only silence — speak louder/closer." Each one names the likely physical cause and the next command. That came from watching real sessions: transcription failures are almost never model failures, they're room failures — muted mics, wrong devices, whisper-quiet laptops. So sk mic-test records three seconds and returns a verdict — silent, quiet, good — with peak dB and an alsamixer hint when it's the OS mixer lying to you.
The privacy side is structural, not promised. Recordings are temp files deleted after every take — there's nowhere for them to accumulate because nothing in the pipeline has a server to send them to. The strongest privacy guarantee is an architecture with no upload path.
What broke
Two favorites from src/sk/voice.py. First: Python 3.13 removed audioop, so mic-level measurement needed a hand-rolled PCM stats function with struct — including a comment about endianness placement in format strings that reads like someone lost an afternoon to it.
Second, the spawn guard. faster-whisper's dependency tree can start a multiprocessing child, and on macOS recycled file descriptors made CPython reject the spawn with ValueError: bad value(s) in fds_to_keep. The fix monkeypatches spawnv_passfds to sanitize the fd list — dedupe, sort, drop closed ones — installed once and only around STT work. And because callers reduce failures to one-line messages, which hid the genuine crash site, full tracebacks now append to ~/.sidekick/voice-errors.log. The bug that teaches you your error reporting is broken is always two bugs.
Part 4 goes into the tool safety model: 19 tools, approval gates, hard refusals, and the audit ledger.
Built by the Sidekick community. Repo: https://github.com/Faisal-Fayaz/sidekick
Top comments (0)