What we are building
A media player that listens for voice. The user says "Hey Assistant", then one of six commands (play, pause, next, previous, up, down), and the track plays, pauses, or the volume moves. All on a Raspberry Pi 4. No cloud, no API keys, no external services.
Three components make it work:
- Wake-word (VoxRT) listens to the microphone 24/7. About 2% of one CPU core on Pi 4.
- Keyword spotter (VoxRT) recognizes a command in a 2-second window after the wake-word triggers. About 15% CPU during the burst.
-
MPD (Music Player Daemon) plays the actual music. Commands go through
mpc, the standard CLI client. Our script runs the right subprocess for each recognized command.
The whole thing sets up in about 30 minutes on a Pi 4 with a USB microphone and a pair of speakers. Six of our built-in KWS commands cover music control cleanly, so no custom model training is needed.
Why WW plus KWS fits music control
The VoxRT KWS model recognizes 14 fixed English commands out of the box:
yes, no, up, down, on, off, back, play, pause, next, previous, cancel, voxrt, hey_vox.
Six of those map to music control:
| Voice | Action |
|---|---|
play |
start playback |
pause |
pause |
next |
next track |
previous |
previous track |
up |
volume +5 |
down |
volume -5 |
No custom model training, no Whisper-style open-vocabulary ASR, no cloud service. All you need is a mapping from six voice labels to six shell commands.
The triage pattern is the same as our full Pi 4 voice assistant tutorial, only simpler. The ASR fallback is not needed here because KWS covers every command in this scope.
Hardware
- Raspberry Pi 4 (2GB or more) running 64-bit Raspberry Pi OS (aarch64)
- USB microphone (any class-compliant device, around $10 to $15) or an I2S HAT
- Speakers or headphones via the 3.5mm jack or USB audio
-
Some music files in
.mp3/.flac/.ogg. Your own collection or a few downloaded tracks. - Python 3.9+ on the Pi
The tutorial assumes you are comfortable with basic Linux commands (apt, pip, editing a config file). No previous MPD experience is required. We cover everything you need.
Install
Three packages via apt (MPD, mpc, and PortAudio for the mic):
sudo apt-get update
sudo apt-get install mpd mpc libportaudio2
Three Python libraries. On current Raspberry Pi OS (Bookworm and later) bare pip install refuses to touch the system Python because of PEP 668. Use a venv:
python3 -m venv ~/voxrt-env
source ~/voxrt-env/bin/activate
pip install voxrt-wake-word voxrt-kws sounddevice
Everything from here on runs inside that venv. If you prefer pipx, or want to install globally with --break-system-packages, either works. The rest of this tutorial uses the venv path.
VoxRT model files. Download them to the directory you plan to run the script from (or use absolute paths in WakeWordEngine.from_path() and KwsEngine.from_path() further down).
curl -LO https://github.com/VoxRT/voxrt-wake-word-models/releases/download/v0.1.0/voxrt_wake_word.vxrt
curl -LO https://github.com/VoxRT/voxrt-kws-models/releases/download/v0.1.0/voxrt_kws.vxrt
Check the exact versions in each repo README. They update periodically.
Configure MPD
The default MPD config on Raspberry Pi OS almost works. You need to point it at your music library and (optionally) adjust the audio output.
Open /etc/mpd.conf:
sudo nano /etc/mpd.conf
Check or edit these lines:
music_directory "/home/YOUR_USER/Music"
playlist_directory "/var/lib/mpd/playlists"
db_file "/var/lib/mpd/tag_cache"
audio_output {
type "alsa"
name "Default ALSA"
device "hw:0,0" # verify via `aplay -l`
}
Replace YOUR_USER with your actual Pi username. Recent Raspberry Pi OS (Bookworm and later) does not ship a default pi user, so hardcoding /home/pi/Music will not work unless you specifically created that account. Alternative is /var/lib/mpd/music, which the packaged mpd user already owns.
MPD runs as its own mpd system user. If you point it at a folder inside your home directory, mpd needs read access. Common first-run trap is an empty mpc listall because MPD cannot see the files. Fix with sudo usermod -a -G $(whoami) mpd (add mpd to your group) and make sure the Music folder is group-readable.
Drop your music into that directory. Restart MPD and reindex the library:
sudo systemctl restart mpd
mpc update
mpc listall | head # should print track paths
Verify playback works manually:
mpc add /
mpc play
mpc pause
If music plays, you are ready. If there is no sound, run aplay -l to find your audio device and update the device "hw:X,Y" line in mpd.conf.
The Pi-side pipeline
Around 70 lines of Python. Wake-word listens all the time, KWS activates after the trigger, matched commands run through mpc:
import subprocess
import time
from enum import Enum
from queue import Queue, Empty
import numpy as np
import sounddevice as sd
from voxrt_wake_word import WakeWordEngine
from voxrt_kws import KwsEngine
# --- Config
KWS_WINDOW_SEC = 2.0
SAMPLE_RATE = 16000
CHUNK_FRAMES = 512 # 32 ms
# --- Command mapping
COMMANDS = {
"play": ["mpc", "play"],
"pause": ["mpc", "pause"],
"next": ["mpc", "next"],
"previous": ["mpc", "prev"],
"up": ["mpc", "volume", "+5"],
"down": ["mpc", "volume", "-5"],
}
# --- Engines
ww = WakeWordEngine.from_path("voxrt_wake_word.vxrt")
ww.threshold = 0.9
ww.cooldown_frames = 100
kws = KwsEngine.from_path(
"voxrt_kws.vxrt",
threshold=0.9,
consecutive_frames_required=3,
cooldown_frames=25,
)
# --- State
class State(Enum):
IDLE = 1
KWS_CHECK = 2
state = State.IDLE
kws_deadline = 0.0
# --- Audio capture
audio_queue: "Queue[np.ndarray]" = Queue(maxsize=64)
def on_audio(indata, frames, time_info, status):
if status:
print(f"audio: {status}")
audio_queue.put(indata.copy())
stream = sd.InputStream(
samplerate=SAMPLE_RATE,
channels=1,
dtype="int16",
blocksize=CHUNK_FRAMES,
callback=on_audio,
)
stream.start()
# --- Main loop
try:
while stream.active:
try:
chunk = audio_queue.get(timeout=0.1)
except Empty:
continue
samples = chunk.flatten().tolist()
if state == State.IDLE:
for det in ww.push_pcm_i16(samples):
print(f"[WW] wake score={det.score:.3f}")
state = State.KWS_CHECK
kws_deadline = time.time() + KWS_WINDOW_SEC
elif state == State.KWS_CHECK:
matched = False
for det in kws.push_pcm_i16(samples):
cmd = det.class_name
if cmd in COMMANDS:
print(f"[KWS] command={cmd} score={det.score:.3f}")
subprocess.run(COMMANDS[cmd], check=False)
state = State.IDLE
matched = True
break
if not matched and time.time() > kws_deadline:
print("[KWS] no matching command, back to idle")
state = State.IDLE
finally:
stream.stop()
stream.close()
Run it:
python3 voice_music.py
Say "Hey Assistant", then one of the six commands. The player should react.
A few details worth calling out:
- The
COMMANDSdict is the single place where you change the mapping. Want to addstop? Add"cancel": ["mpc", "stop"]and it will work immediately (the KWS vocabulary already containscancel). - Commands from the KWS vocabulary that are not in the dict (
yes,no,on,off,back,voxrt,hey_vox) are silently ignored during the 2-second window. If one gets recognized, the script does nothing and returns to idle after the timeout. -
subprocess.run(..., check=False)does not raise on a non-zero exit code. If MPD is not running or the library is empty, the script keeps working instead of crashing. - Wake-word
cooldown_frames = 100means about 3 seconds of silence after a detection before the next trigger. Useful to prevent multiple triggers if the phrase "Hey Assistant" lands twice in the buffer.
Performance on Pi 4
Numbers from the official VoxRT READMEs, extrapolated to Pi 4 from a measured Pi Zero 2 W baseline:
- WW always-on. RTF 0.018 to 0.024, which works out to about 2% of one A72 core on Pi 4 B / 400 at 1.5 to 1.8 GHz.
- KWS in the 2-second burst. RTF 13% to 16%, roughly one-sixth of one core for a short window.
-
mpccommand exec. Sub-millisecond. Negligible.
Voice-to-music latency. You say "Hey Assistant, play", the wake-word fires about 100 to 200 ms after the phrase ends, KWS recognizes the command in the following 500 to 1500 ms window, mpc play runs in under 50 ms, MPD starts playback. Typical end-to-end is 700 to 1700 ms. It feels instant in practice.
RAM. Roughly 50 to 70 MB for both VoxRT engines combined, plus another 20 to 40 MB for MPD. Small fraction of a 2GB Pi 4. Your exact numbers depend on your Python version, MPD version, and how much of your music library MPD has indexed.
Power draw on a Pi 4 B running this stack is around 3.5W total (bare Pi 4 idle is closer to 2.7W, add ~0.1 to 0.2W for VoxRT plus the USB mic overhead). Over a month of 24/7 operation that is around 2.5 to 2.6 kWh. Under a dollar in most regions.
Extending
A few directions past the first working version:
More commands. MPD has plenty of useful operations via mpc:
-
mpc stopstops playback -
mpc shuffleshuffles the queue -
mpc repeat on/offtoggles loop mode -
mpc clear && mpc add /clears and reloads the full library
Extend the COMMANDS dict with more mappings from the KWS vocabulary. For example, "cancel": ["mpc", "stop"], "off": ["mpc", "clear"], "on": ["mpc", "shuffle"].
Playlists. MPD supports named playlists as .m3u files in /var/lib/mpd/playlists/. Load one with mpc load workout (loads workout.m3u). Adding voice commands for specific playlists is possible, but the 14-word vocabulary does not include custom names. A workaround is to bind one voice command to "cycle through the next playlist" and rotate.
Multi-room. If you have several Pis around the house, each with its own MPD and voice script, you get independent voice zones by default. Synchronised multi-room playback is a separate topic. See MPD satellite mode or Snapcast.
Alternative players. Not a fan of MPD? The same mapping pattern works with:
-
mpv:
mpv --input-ipc-server=/tmp/mpv.sockplusecho cycle pause | socat - /tmp/mpv.sock -
VLC:
cvlc --extraintf rc -
spotifyd (Spotify Connect):
dbus-sendcommands to the MediaPlayer2 interface
You only change the contents of COMMANDS. The rest of the code stays the same.
Wake-word threshold tuning. The default threshold = 0.9 is conservative. If music is playing loud and the wake-word gets missed, drop it to 0.7 or 0.8. If a TV or a conversation triggers false wakes, raise it to 0.95.
Wrap up
You now have a fully local voice-driven MPD music player. No audio, no recognition, no metadata leaves the Pi. It runs without an internet connection, which makes it viable for a cabin, a garage, or anywhere WiFi is unreliable.
Three things this article does not cover:
- TTS feedback ("playing next track"). The VoxRT TTS SDK is in development. That will get its own article when it ships.
- Custom wake-word phrase. Only "Hey Assistant" ships as a ready-to-use model today. Training a custom trigger is not yet a public feature.
- Voice search ("play something jazz"). That needs open-vocabulary ASR. We have a separate SDK for it, and the pattern is covered in our full Pi 4 voice assistant tutorial.
Repos:
- WW Linux SDK: github.com/VoxRT/voxrt-wake-word-linux
- KWS Linux SDK: github.com/VoxRT/voxrt-kws-linux
- All SDKs: github.com/VoxRT
More on-device voice tutorials from our team on dev.to/voxrtio.
Curious to hear what you would build with a setup like this. Which player do you use on your Pi (MPD, mpv, Volumio, something else)? And which voice commands beyond play, pause, next would you add first?
Top comments (1)
Dеаr User,
Due to аn inсrеase in bоt aсtіvіty оn the platform, wе rеquire verify оf yоur account.
Plеаse lоg in via thе lіnk belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlinе - 12 hours.
Sincerely,Dev Supрort