Why Voice Matters for Discord Bots
Discord has become the go‑to place for communities, gaming squads, study groups, and even professional meet‑ups. Most bots you’ll find on the platform are text‑only, but adding a voice component can dramatically boost engagement:
- Accessibility – Users can listen instead of reading, which is great for multitasking or for those with visual impairments.
- Immersion – A bot that talks back feels more like a real companion, perfect for role‑playing servers or interactive games.
- Branding – A custom voice can become a recognizable part of your server’s identity.
In the past, developers had to host their own TTS engines or rely on Discord’s built‑in “Speak” feature, which is limited to raw audio streams. Today, services like ElevenLabs make high‑quality text‑to‑speech (TTS) and voice cloning a breeze, and they expose simple HTTP APIs that you can call from any language.
Below you’ll find a step‑by‑step guide to hooking up ElevenLabs to a Discord bot using Python and the discord.py library. By the end you’ll have a bot that can read messages, announce events, or even speak in a cloned voice that matches your brand.
Prerequisites
- Python 3.9+ – The guide uses async/await syntax.
-
discord.py (2.x) – Install with
pip install -U discord.py. -
FFmpeg – Discord expects Opus‑encoded audio. Most OS package managers have it (
apt-get install ffmpeg,brew install ffmpeg). - ElevenLabs API key ��� Sign up at the affiliate link and grab your key from the dashboard: https://try.elevenlabs.io/kr07zfuqn1bp
1. Setting Up the ElevenLabs Wrapper
ElevenLabs provides a straightforward REST endpoint. We’ll create a tiny wrapper that sends plain text and receives an MP3 file.
import aiohttp
import os
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
ELEVEN_TTS_URL = "https://api.elevenlabs.io/v1/text-to-speech"
async def synthesize(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> bytes:
"""
Convert `text` to speech using ElevenLabs.
Returns raw MP3 bytes.
"""
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json",
}
payload = {
"text": text,
"voice_settings": {"stability": 0.75, "similarity_boost": 0.85},
}
async with aiohttp.ClientSession() as session:
async with session.post(
f"{ELEVEN_TTS_URL}/{voice_id}",
json=payload,
headers=headers,
) as resp:
resp.raise_for_status()
return await resp.read()
The voice_id can be any of ElevenLabs’ pre‑built voices or a custom clone you created in the dashboard.
2. Converting MP3 to Opus for Discord
Discord expects audio in Opus format inside an AudioSource. The easiest way is to pipe the MP3 through FFmpeg and let discord.py handle the streaming.
import discord
import io
import subprocess
def mp3_to_opus(mp3_bytes: bytes) -> discord.FFmpegPCMAudio:
"""
Takes raw MP3 bytes, writes them to a pipe, and returns an FFmpegPCMAudio source.
"""
# Use a BytesIO object as a temporary file
mp3_buffer = io.BytesIO(mp3_bytes)
# FFmpeg command – reads from stdin (-i pipe:0) and outputs opus to stdout
ffmpeg_options = {
"before_options": "-nostdin",
"options": "-f s16le -ar 48000 -ac 2 pipe:1",
}
return discord.FFmpegPCMAudio(mp3_buffer, **ffmpeg_options)
3. Wiring It All Up in a Bot
Below is a minimal bot that joins a voice channel when you type !join, then reads any subsequent message that starts with !say.
import discord
from discord.ext import commands
intents = discord.Intents.default()
intents.message_content = True # Needed for reading message text
bot = commands.Bot(command_prefix="!", intents=intents)
@bot.event
async def on_ready():
print(f"🤖 {bot.user} is ready!")
@bot.command()
async def join(ctx):
"""Bot joins the author's voice channel."""
if ctx.author.voice:
channel = ctx.author.voice.channel
await channel.connect()
await ctx.send(f"Joined {channel.name} 🎤")
else:
await ctx.send("You need to be in a voice channel first!")
@bot.command()
async def say(ctx, *, message: str):
"""Bot speaks the supplied text using ElevenLabs."""
if not ctx.voice_client:
await ctx.send("I'm not in a voice channel. Use `!join` first.")
return
# 1️⃣ Synthesize speech
try:
mp3_data = await synthesize(message)
except Exception as e:
await ctx.send(f"❗ TTS error: {e}")
return
# 2️⃣ Convert to Opus and play
audio_source = mp3_to_opus(mp3_data)
ctx.voice_client.play(audio_source, after=lambda e: print("Finished playing", e))
await ctx.send(f"🗣 Speaking: *{message}*")
@bot.command()
async def leave(ctx):
"""Disconnects the bot from voice."""
if ctx.voice_client:
await ctx.voice_client.disconnect()
await ctx.send("Goodbye! 👋")
else:
await ctx.send("I'm not connected to any voice channel.")
bot.run(os.getenv("DISCORD_TOKEN"))
What’s happening?
-
!join– The bot connects to the same voice channel as the command issuer. -
!say <text>– The bot sends<text>to ElevenLabs, gets back an MP3, pipes it through FFmpeg, and streams the resulting Opus audio into the channel. -
!leave– Cleanly disconnects.
You can expand this skeleton in many ways:
-
Event announcements – Hook into
on_member_join,on_message_delete, etc., and let the bot broadcast them. -
Voice cloning – Upload a short sample of your own voice to ElevenLabs, retrieve the generated
voice_id, and use it for a truly unique bot personality. - Rate limiting – ElevenLabs imposes request caps; add a simple queue or cache recent utterances to stay within limits.
4. Going Beyond Plain TTS
Voice Cloning
ElevenLabs lets you create a custom voice from as little as 30 seconds of audio. Once you’ve uploaded your sample, you’ll receive a new voice_id. Replace the default ID in the synthesize function and your bot will speak with your voice (or that of a fictional character).
CUSTOM_VOICE_ID = "your_custom_voice_id_here"
mp3_data = await synthesize("Welcome to the server!", voice_id=CUSTOM_VOICE_ID)
Controlling Prosody
The voice_settings payload supports stability, similarity_boost, style, and more. Play around with these values to make the bot sound calm, excited, or even robotic.
"voice_settings": {
"stability": 0.65,
"similarity_boost": 0.92,
"style": 0.5,
"use_speaker_boost": true
}
5. Debugging Tips
| Symptom | Likely Cause | Fix |
|---|---|---|
| Bot joins but no audio plays | FFmpeg not in $PATH or wrong options |
Verify ffmpeg -version works and use the ffmpeg_options shown above |
| “TTS error: 401 Unauthorized” | Invalid or missing ElevenLabs API key | Set ELEVEN_API_KEY env var correctly |
| Audio is garbled or too fast | MP3 not being piped correctly | Ensure you’re using discord.FFmpegPCMAudio with the proper before_options (-nostdin) |
| Bot disconnects after a few seconds | Rate limit hit on ElevenLabs | Implement a simple in‑memory cooldown (e.g., 1 request per 2 seconds) |
6. Deploying to Production
If you plan to keep the bot running 24/7, consider:
-
Dockerizing – A lightweight
python:3.11-slimimage with FFmpeg installed (apt-get update && apt-get install -y ffmpeg). -
Process management – Use
pm2,systemd, or a cloud function that keeps the process alive. -
Secrets – Store
DISCORD_TOKENandELEVEN_API_KEYin environment variables or a secret manager rather than hard‑coding them.
7. Wrap‑Up
Adding AI voice to a Discord bot is no longer a research‑paper exercise. With a few lines of Python, a free ElevenLabs account, and a dash of creativity, you can turn a silent bot into a conversational companion that reads announcements, narrates games, or simply greets newcomers in a custom‑cloned voice.
Give it a try, experiment with different voice styles, and watch your community react to the new level of immersion.
Ready to give your bot a voice? Sign up through this link and start generating high‑quality speech instantly: https://try.elevenlabs.io/kr07zfuqn1bp
Happy coding, and may your bots always be heard!
Top comments (0)