Introduction
Imagine a chatbot that can not only answer questions but also talk like a whole cast of characters—think a narrator, a quirky side‑kick, and a stern instructor—all in the same conversation. With modern text‑to‑speech (TTS) services you can generate lifelike speech on the fly, and with a little glue code you can stitch those snippets together into a seamless multi‑voice dialogue.
In this article we’ll build a Multi‑Voice Dialogue Generator from scratch. You’ll learn how to:
- Set up the ElevenLabs TTS API (our recommended voice‑cloning platform).
- Define multiple characters, each with its own voice settings.
- Generate audio for each line of a scripted dialogue.
- Concatenate the audio files into a single playback stream.
All of the code is in Python, but the same concepts translate to JavaScript, curl, or any language that can make HTTP requests.
Prerequisites
| What you need | Why |
|---|---|
| Python 3.8+ | The demo script uses requests and pydub. |
| pip | To install the required libraries. |
| ElevenLabs API key | Enables you to call the TTS endpoint. Sign up at the official site and grab a key. |
| ffmpeg (system‑wide) |
pydub relies on it for audio concatenation. |
# Install Python deps
pip install requests pydub tqdm
# On macOS (brew) or Ubuntu (apt)
brew install ffmpeg # macOS
sudo apt-get install ffmpeg # Ubuntu
Setting up the ElevenLabs API
ElevenLabs offers a powerful REST API that can synthesize speech in a variety of voices, including custom clones. Sign up and head to the API Keys section of the dashboard. Once you have your secret token, keep it safe; you’ll need it for every request.
import os
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY") # export ELEVEN_API_KEY=your_key
BASE_URL = "https://api.elevenlabs.io/v1"
HEADERS = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
If you don’t have an API key yet, you can get one instantly by signing up through this link: ElevenLabs. The affiliate link gives you a quick start with a free credit bundle.
Defining Your Characters
Each character in the dialogue will be mapped to a specific voice ID from ElevenLabs. The platform provides a set of pre‑built voices (e.g., “Rachel”, “Domi”) and lets you upload custom voice clones. Grab the IDs from the Voices tab in the dashboard.
# Example character map
CHARACTERS = {
"Narrator": "EXAVITQu4vr4xnSDxMaL", # pre‑built voice ID
"Sidekick": "V1ZJjVvV0KZyK6x5LZVZ", # custom clone
"Instructor": "Tx8E9A6sG2VhW2W8JkzU" # another pre‑built voice
}
Feel free to add more entries or change the voice IDs to match your own clones. The key point is that each line of dialogue will be routed to the appropriate voice.
Generating Audio for Each Line
Below is a helper function that sends a single line of text to ElevenLabs and saves the resulting MP3 to disk. We’ll also use tqdm for a nice progress bar when processing many lines.
import requests
from pathlib import Path
from tqdm import tqdm
def synthesize_line(text: str, voice_id: str, outfile: Path):
"""Call ElevenLabs TTS and write the audio to outfile."""
url = f"{BASE_URL}/text-to-speech/{voice_id}"
payload = {
"text": text,
"model_id": "eleven_monolingual_v1", # default high‑quality model
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
response = requests.post(url, json=payload, headers=HEADERS, stream=True)
response.raise_for_status()
with open(outfile, "wb") as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
# Example usage
line = "Welcome to the adventure, brave explorer!"
synthesize_line(line, CHARACTERS["Narrator"], Path("output/01_narrator.mp3"))
If you prefer a quick curl test, here’s the equivalent command:
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \
-H "xi-api-key: $ELEVEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Hello, world!",
"model_id": "eleven_monolingual_v1",
"voice_settings": {"stability":0.75,"similarity_boost":0.85}
}' --output hello.mp3
Stitching the Pieces Together
Once each line is rendered, we need to concatenate them in the correct order. pydub makes this trivial.
from pydub import AudioSegment
def concatenate_clips(folder: Path, ordered_files: list[Path], out_file: Path):
"""Merge a list of MP3 files into a single audio track."""
combined = AudioSegment.empty()
for fp in ordered_files:
clip = AudioSegment.from_file(fp)
combined += clip
combined.export(out_file, format="mp3")
# Build a simple script
script = [
("Narrator", "Welcome to the adventure, brave explorer!"),
("Sidekick", "Whoa, did you just see that dragon?"),
("Instructor", "Remember, keep your shield up at all times."),
("Sidekick", "Got it! Let's roll!"),
]
# Create output dir
output_dir = Path("output")
output_dir.mkdir(exist_ok=True)
# Generate each line
audio_files = []
for i, (speaker, line) in enumerate(script, start=1):
voice_id = CHARACTERS[speaker]
out_path = output_dir / f"{i:02d}_{speaker.lower()}.mp3"
synthesize_line(line, voice_id, out_path)
audio_files.append(out_path)
# Concatenate
final_path = output_dir / "dialogue.mp3"
concatenate_clips(output_dir, audio_files, final_path)
print(f"✅ Dialogue saved to {final_path}")
Run the script and you’ll get a single dialogue.mp3 that sounds like a mini‑audio drama, with each character speaking in its own distinct voice.
Testing and Tweaking
Adjusting Voice Settings
ElevenLabs lets you fine‑tune two main parameters:
| Setting | Effect |
|---|---|
| stability | Controls how consistent the voice sounds across sentences. Higher values reduce wobble. |
| similarity_boost | Pushes the output closer to the reference voice (useful for clones). |
Play with values between 0.0 and 1.0 to find the sweet spot for your characters.
Adding Pauses
Natural conversation includes brief silences. You can inject a short silent clip between lines:
silence = AudioSegment.silent(duration=300) # 300 ms
combined = AudioSegment.empty()
for clip_path in audio_files:
combined += AudioSegment.from_file(clip_path)
combined += silence
combined.export(final_path, format="mp3")
Streaming Directly to a Browser
If you want to serve the dialogue on a web page, simply expose the final MP3 via an endpoint (Flask, FastAPI, etc.) and use an <audio> tag:
<audio controls src="/static/dialogue.mp3"></audio>
Wrap‑up
You now have a fully functional pipeline that:
- Maps characters to distinct ElevenLabs voices
- Synthesizes each line on demand
- Combines the snippets into a coherent audio file
The same approach works for interactive bots, game NPCs, or even automated podcasts where multiple hosts converse. The key is keeping the voice IDs and script data separate so you can swap characters in and out without touching the synthesis logic.
Call to Action
Ready to give your applications a real‑world voice? Sign up for ElevenLabs today through this link: ElevenLabs. Grab your API key, spin up the script above, and start building immersive multi‑voice experiences right away. Happy coding!
Top comments (0)