DEV Community

Cover image for Build a voice-controlled music player on Raspberry Pi 4 (fully offline)
VoxRT
VoxRT

Posted on

Build a voice-controlled music player on Raspberry Pi 4 (fully offline)

What we are building

A media player that listens for voice. The user says "Hey Assistant", then one of six commands (play, pause, next, previous, up, down), and the track plays, pauses, or the volume moves. All on a Raspberry Pi 4. No cloud, no API keys, no external services.

Three components make it work:

  • Wake-word (VoxRT) listens to the microphone 24/7. About 2% of one CPU core on Pi 4.
  • Keyword spotter (VoxRT) recognizes a command in a 2-second window after the wake-word triggers. About 15% CPU during the burst.
  • MPD (Music Player Daemon) plays the actual music. Commands go through mpc, the standard CLI client. Our script runs the right subprocess for each recognized command.

The whole thing sets up in about 30 minutes on a Pi 4 with a USB microphone and a pair of speakers. Six of our built-in KWS commands cover music control cleanly, so no custom model training is needed.

Why WW plus KWS fits music control

The VoxRT KWS model recognizes 14 fixed English commands out of the box:
yes, no, up, down, on, off, back, play, pause, next, previous, cancel, voxrt, hey_vox.

Six of those map to music control:

Voice Action
play start playback
pause pause
next next track
previous previous track
up volume +5
down volume -5

No custom model training, no Whisper-style open-vocabulary ASR, no cloud service. All you need is a mapping from six voice labels to six shell commands.

The triage pattern is the same as our full Pi 4 voice assistant tutorial, only simpler. The ASR fallback is not needed here because KWS covers every command in this scope.

Hardware

  • Raspberry Pi 4 (2GB or more) running 64-bit Raspberry Pi OS (aarch64)
  • USB microphone (any class-compliant device, around $10 to $15) or an I2S HAT
  • Speakers or headphones via the 3.5mm jack or USB audio
  • Some music files in .mp3 / .flac / .ogg. Your own collection or a few downloaded tracks.
  • Python 3.9+ on the Pi

The tutorial assumes you are comfortable with basic Linux commands (apt, pip, editing a config file). No previous MPD experience is required. We cover everything you need.

Install

Three packages via apt (MPD, mpc, and PortAudio for the mic):

sudo apt-get update
sudo apt-get install mpd mpc libportaudio2
Enter fullscreen mode Exit fullscreen mode

Three Python libraries. On current Raspberry Pi OS (Bookworm and later) bare pip install refuses to touch the system Python because of PEP 668. Use a venv:

python3 -m venv ~/voxrt-env
source ~/voxrt-env/bin/activate
pip install voxrt-wake-word voxrt-kws sounddevice
Enter fullscreen mode Exit fullscreen mode

Everything from here on runs inside that venv. If you prefer pipx, or want to install globally with --break-system-packages, either works. The rest of this tutorial uses the venv path.

VoxRT model files. Download them to the directory you plan to run the script from (or use absolute paths in WakeWordEngine.from_path() and KwsEngine.from_path() further down).

curl -LO https://github.com/VoxRT/voxrt-wake-word-models/releases/download/v0.1.0/voxrt_wake_word.vxrt
curl -LO https://github.com/VoxRT/voxrt-kws-models/releases/download/v0.1.0/voxrt_kws.vxrt
Enter fullscreen mode Exit fullscreen mode

Check the exact versions in each repo README. They update periodically.

Configure MPD

The default MPD config on Raspberry Pi OS almost works. You need to point it at your music library and (optionally) adjust the audio output.

Open /etc/mpd.conf:

sudo nano /etc/mpd.conf
Enter fullscreen mode Exit fullscreen mode

Check or edit these lines:

music_directory     "/home/YOUR_USER/Music"
playlist_directory  "/var/lib/mpd/playlists"
db_file             "/var/lib/mpd/tag_cache"

audio_output {
    type            "alsa"
    name            "Default ALSA"
    device          "hw:0,0"   # verify via `aplay -l`
}
Enter fullscreen mode Exit fullscreen mode

Replace YOUR_USER with your actual Pi username. Recent Raspberry Pi OS (Bookworm and later) does not ship a default pi user, so hardcoding /home/pi/Music will not work unless you specifically created that account. Alternative is /var/lib/mpd/music, which the packaged mpd user already owns.

MPD runs as its own mpd system user. If you point it at a folder inside your home directory, mpd needs read access. Common first-run trap is an empty mpc listall because MPD cannot see the files. Fix with sudo usermod -a -G $(whoami) mpd (add mpd to your group) and make sure the Music folder is group-readable.

Drop your music into that directory. Restart MPD and reindex the library:

sudo systemctl restart mpd
mpc update
mpc listall | head    # should print track paths
Enter fullscreen mode Exit fullscreen mode

Verify playback works manually:

mpc add /
mpc play
mpc pause
Enter fullscreen mode Exit fullscreen mode

If music plays, you are ready. If there is no sound, run aplay -l to find your audio device and update the device "hw:X,Y" line in mpd.conf.

The Pi-side pipeline

Around 70 lines of Python. Wake-word listens all the time, KWS activates after the trigger, matched commands run through mpc:

import subprocess
import time
from enum import Enum
from queue import Queue, Empty

import numpy as np
import sounddevice as sd

from voxrt_wake_word import WakeWordEngine
from voxrt_kws import KwsEngine

# --- Config
KWS_WINDOW_SEC = 2.0
SAMPLE_RATE = 16000
CHUNK_FRAMES = 512   # 32 ms

# --- Command mapping
COMMANDS = {
    "play":     ["mpc", "play"],
    "pause":    ["mpc", "pause"],
    "next":     ["mpc", "next"],
    "previous": ["mpc", "prev"],
    "up":       ["mpc", "volume", "+5"],
    "down":     ["mpc", "volume", "-5"],
}

# --- Engines
ww = WakeWordEngine.from_path("voxrt_wake_word.vxrt")
ww.threshold = 0.9
ww.cooldown_frames = 100

kws = KwsEngine.from_path(
    "voxrt_kws.vxrt",
    threshold=0.9,
    consecutive_frames_required=3,
    cooldown_frames=25,
)

# --- State
class State(Enum):
    IDLE = 1
    KWS_CHECK = 2

state = State.IDLE
kws_deadline = 0.0

# --- Audio capture
audio_queue: "Queue[np.ndarray]" = Queue(maxsize=64)

def on_audio(indata, frames, time_info, status):
    if status:
        print(f"audio: {status}")
    audio_queue.put(indata.copy())

stream = sd.InputStream(
    samplerate=SAMPLE_RATE,
    channels=1,
    dtype="int16",
    blocksize=CHUNK_FRAMES,
    callback=on_audio,
)
stream.start()

# --- Main loop
try:
    while stream.active:
        try:
            chunk = audio_queue.get(timeout=0.1)
        except Empty:
            continue

        samples = chunk.flatten().tolist()

        if state == State.IDLE:
            for det in ww.push_pcm_i16(samples):
                print(f"[WW] wake score={det.score:.3f}")
                state = State.KWS_CHECK
                kws_deadline = time.time() + KWS_WINDOW_SEC

        elif state == State.KWS_CHECK:
            matched = False
            for det in kws.push_pcm_i16(samples):
                cmd = det.class_name
                if cmd in COMMANDS:
                    print(f"[KWS] command={cmd} score={det.score:.3f}")
                    subprocess.run(COMMANDS[cmd], check=False)
                    state = State.IDLE
                    matched = True
                    break

            if not matched and time.time() > kws_deadline:
                print("[KWS] no matching command, back to idle")
                state = State.IDLE
finally:
    stream.stop()
    stream.close()
Enter fullscreen mode Exit fullscreen mode

Run it:

python3 voice_music.py
Enter fullscreen mode Exit fullscreen mode

Say "Hey Assistant", then one of the six commands. The player should react.

A few details worth calling out:

  • The COMMANDS dict is the single place where you change the mapping. Want to add stop? Add "cancel": ["mpc", "stop"] and it will work immediately (the KWS vocabulary already contains cancel).
  • Commands from the KWS vocabulary that are not in the dict (yes, no, on, off, back, voxrt, hey_vox) are silently ignored during the 2-second window. If one gets recognized, the script does nothing and returns to idle after the timeout.
  • subprocess.run(..., check=False) does not raise on a non-zero exit code. If MPD is not running or the library is empty, the script keeps working instead of crashing.
  • Wake-word cooldown_frames = 100 means about 3 seconds of silence after a detection before the next trigger. Useful to prevent multiple triggers if the phrase "Hey Assistant" lands twice in the buffer.

Performance on Pi 4

Numbers from the official VoxRT READMEs, extrapolated to Pi 4 from a measured Pi Zero 2 W baseline:

  • WW always-on. RTF 0.018 to 0.024, which works out to about 2% of one A72 core on Pi 4 B / 400 at 1.5 to 1.8 GHz.
  • KWS in the 2-second burst. RTF 13% to 16%, roughly one-sixth of one core for a short window.
  • mpc command exec. Sub-millisecond. Negligible.

Voice-to-music latency. You say "Hey Assistant, play", the wake-word fires about 100 to 200 ms after the phrase ends, KWS recognizes the command in the following 500 to 1500 ms window, mpc play runs in under 50 ms, MPD starts playback. Typical end-to-end is 700 to 1700 ms. It feels instant in practice.

RAM. Roughly 50 to 70 MB for both VoxRT engines combined, plus another 20 to 40 MB for MPD. Small fraction of a 2GB Pi 4. Your exact numbers depend on your Python version, MPD version, and how much of your music library MPD has indexed.

Power draw on a Pi 4 B running this stack is around 3.5W total (bare Pi 4 idle is closer to 2.7W, add ~0.1 to 0.2W for VoxRT plus the USB mic overhead). Over a month of 24/7 operation that is around 2.5 to 2.6 kWh. Under a dollar in most regions.

Extending

A few directions past the first working version:

More commands. MPD has plenty of useful operations via mpc:

  • mpc stop stops playback
  • mpc shuffle shuffles the queue
  • mpc repeat on/off toggles loop mode
  • mpc clear && mpc add / clears and reloads the full library

Extend the COMMANDS dict with more mappings from the KWS vocabulary. For example, "cancel": ["mpc", "stop"], "off": ["mpc", "clear"], "on": ["mpc", "shuffle"].

Playlists. MPD supports named playlists as .m3u files in /var/lib/mpd/playlists/. Load one with mpc load workout (loads workout.m3u). Adding voice commands for specific playlists is possible, but the 14-word vocabulary does not include custom names. A workaround is to bind one voice command to "cycle through the next playlist" and rotate.

Multi-room. If you have several Pis around the house, each with its own MPD and voice script, you get independent voice zones by default. Synchronised multi-room playback is a separate topic. See MPD satellite mode or Snapcast.

Alternative players. Not a fan of MPD? The same mapping pattern works with:

  • mpv: mpv --input-ipc-server=/tmp/mpv.sock plus echo cycle pause | socat - /tmp/mpv.sock
  • VLC: cvlc --extraintf rc
  • spotifyd (Spotify Connect): dbus-send commands to the MediaPlayer2 interface

You only change the contents of COMMANDS. The rest of the code stays the same.

Wake-word threshold tuning. The default threshold = 0.9 is conservative. If music is playing loud and the wake-word gets missed, drop it to 0.7 or 0.8. If a TV or a conversation triggers false wakes, raise it to 0.95.

Wrap up

You now have a fully local voice-driven MPD music player. No audio, no recognition, no metadata leaves the Pi. It runs without an internet connection, which makes it viable for a cabin, a garage, or anywhere WiFi is unreliable.

Three things this article does not cover:

  • TTS feedback ("playing next track"). The VoxRT TTS SDK is in development. That will get its own article when it ships.
  • Custom wake-word phrase. Only "Hey Assistant" ships as a ready-to-use model today. Training a custom trigger is not yet a public feature.
  • Voice search ("play something jazz"). That needs open-vocabulary ASR. We have a separate SDK for it, and the pattern is covered in our full Pi 4 voice assistant tutorial.

Repos:

More on-device voice tutorials from our team on dev.to/voxrtio.

Curious to hear what you would build with a setup like this. Which player do you use on your Pi (MPD, mpv, Volumio, something else)? And which voice commands beyond play, pause, next would you add first?

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеаr User,
Due to аn inсrеase in bоt aсtіvіty оn the platform, wе rеquire verify оf yоur account.
Plеаse lоg in via thе lіnk belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlinе - 12 hours.
Sincerely,Dev Supрort

​‌