DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Build a Multi-Voice Dialogue Generator

Introduction

Imagine a chatbot that can not only answer questions but also talk like a whole cast of characters—think a narrator, a quirky side‑kick, and a stern instructor—all in the same conversation. With modern text‑to‑speech (TTS) services you can generate lifelike speech on the fly, and with a little glue code you can stitch those snippets together into a seamless multi‑voice dialogue.

In this article we’ll build a Multi‑Voice Dialogue Generator from scratch. You’ll learn how to:

  1. Set up the ElevenLabs TTS API (our recommended voice‑cloning platform).
  2. Define multiple characters, each with its own voice settings.
  3. Generate audio for each line of a scripted dialogue.
  4. Concatenate the audio files into a single playback stream.

All of the code is in Python, but the same concepts translate to JavaScript, curl, or any language that can make HTTP requests.


Prerequisites

What you need Why
Python 3.8+ The demo script uses requests and pydub.
pip To install the required libraries.
ElevenLabs API key Enables you to call the TTS endpoint. Sign up at the official site and grab a key.
ffmpeg (system‑wide) pydub relies on it for audio concatenation.
# Install Python deps
pip install requests pydub tqdm

# On macOS (brew) or Ubuntu (apt)
brew install ffmpeg   # macOS
sudo apt-get install ffmpeg   # Ubuntu
Enter fullscreen mode Exit fullscreen mode

Setting up the ElevenLabs API

ElevenLabs offers a powerful REST API that can synthesize speech in a variety of voices, including custom clones. Sign up and head to the API Keys section of the dashboard. Once you have your secret token, keep it safe; you’ll need it for every request.

import os

ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")  # export ELEVEN_API_KEY=your_key
BASE_URL = "https://api.elevenlabs.io/v1"
HEADERS = {
    "xi-api-key": ELEVEN_API_KEY,
    "Content-Type": "application/json"
}
Enter fullscreen mode Exit fullscreen mode

If you don’t have an API key yet, you can get one instantly by signing up through this link: ElevenLabs. The affiliate link gives you a quick start with a free credit bundle.


Defining Your Characters

Each character in the dialogue will be mapped to a specific voice ID from ElevenLabs. The platform provides a set of pre‑built voices (e.g., “Rachel”, “Domi”) and lets you upload custom voice clones. Grab the IDs from the Voices tab in the dashboard.

# Example character map
CHARACTERS = {
    "Narrator": "EXAVITQu4vr4xnSDxMaL",   # pre‑built voice ID
    "Sidekick": "V1ZJjVvV0KZyK6x5LZVZ",   # custom clone
    "Instructor": "Tx8E9A6sG2VhW2W8JkzU" # another pre‑built voice
}
Enter fullscreen mode Exit fullscreen mode

Feel free to add more entries or change the voice IDs to match your own clones. The key point is that each line of dialogue will be routed to the appropriate voice.


Generating Audio for Each Line

Below is a helper function that sends a single line of text to ElevenLabs and saves the resulting MP3 to disk. We’ll also use tqdm for a nice progress bar when processing many lines.

import requests
from pathlib import Path
from tqdm import tqdm

def synthesize_line(text: str, voice_id: str, outfile: Path):
    """Call ElevenLabs TTS and write the audio to outfile."""
    url = f"{BASE_URL}/text-to-speech/{voice_id}"
    payload = {
        "text": text,
        "model_id": "eleven_monolingual_v1",  # default high‑quality model
        "voice_settings": {
            "stability": 0.75,
            "similarity_boost": 0.85
        }
    }

    response = requests.post(url, json=payload, headers=HEADERS, stream=True)
    response.raise_for_status()

    with open(outfile, "wb") as f:
        for chunk in response.iter_content(chunk_size=8192):
            f.write(chunk)

# Example usage
line = "Welcome to the adventure, brave explorer!"
synthesize_line(line, CHARACTERS["Narrator"], Path("output/01_narrator.mp3"))
Enter fullscreen mode Exit fullscreen mode

If you prefer a quick curl test, here’s the equivalent command:

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \
-H "xi-api-key: $ELEVEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "text": "Hello, world!",
  "model_id": "eleven_monolingual_v1",
  "voice_settings": {"stability":0.75,"similarity_boost":0.85}
}' --output hello.mp3
Enter fullscreen mode Exit fullscreen mode

Stitching the Pieces Together

Once each line is rendered, we need to concatenate them in the correct order. pydub makes this trivial.

from pydub import AudioSegment

def concatenate_clips(folder: Path, ordered_files: list[Path], out_file: Path):
    """Merge a list of MP3 files into a single audio track."""
    combined = AudioSegment.empty()
    for fp in ordered_files:
        clip = AudioSegment.from_file(fp)
        combined += clip
    combined.export(out_file, format="mp3")

# Build a simple script
script = [
    ("Narrator", "Welcome to the adventure, brave explorer!"),
    ("Sidekick", "Whoa, did you just see that dragon?"),
    ("Instructor", "Remember, keep your shield up at all times."),
    ("Sidekick", "Got it! Let's roll!"),
]

# Create output dir
output_dir = Path("output")
output_dir.mkdir(exist_ok=True)

# Generate each line
audio_files = []
for i, (speaker, line) in enumerate(script, start=1):
    voice_id = CHARACTERS[speaker]
    out_path = output_dir / f"{i:02d}_{speaker.lower()}.mp3"
    synthesize_line(line, voice_id, out_path)
    audio_files.append(out_path)

# Concatenate
final_path = output_dir / "dialogue.mp3"
concatenate_clips(output_dir, audio_files, final_path)
print(f"✅ Dialogue saved to {final_path}")
Enter fullscreen mode Exit fullscreen mode

Run the script and you’ll get a single dialogue.mp3 that sounds like a mini‑audio drama, with each character speaking in its own distinct voice.


Testing and Tweaking

Adjusting Voice Settings

ElevenLabs lets you fine‑tune two main parameters:

Setting Effect
stability Controls how consistent the voice sounds across sentences. Higher values reduce wobble.
similarity_boost Pushes the output closer to the reference voice (useful for clones).

Play with values between 0.0 and 1.0 to find the sweet spot for your characters.

Adding Pauses

Natural conversation includes brief silences. You can inject a short silent clip between lines:

silence = AudioSegment.silent(duration=300)  # 300 ms
combined = AudioSegment.empty()
for clip_path in audio_files:
    combined += AudioSegment.from_file(clip_path)
    combined += silence
combined.export(final_path, format="mp3")
Enter fullscreen mode Exit fullscreen mode

Streaming Directly to a Browser

If you want to serve the dialogue on a web page, simply expose the final MP3 via an endpoint (Flask, FastAPI, etc.) and use an <audio> tag:

<audio controls src="/static/dialogue.mp3"></audio>
Enter fullscreen mode Exit fullscreen mode

Wrap‑up

You now have a fully functional pipeline that:

  1. Maps characters to distinct ElevenLabs voices
  2. Synthesizes each line on demand
  3. Combines the snippets into a coherent audio file

The same approach works for interactive bots, game NPCs, or even automated podcasts where multiple hosts converse. The key is keeping the voice IDs and script data separate so you can swap characters in and out without touching the synthesis logic.


Call to Action

Ready to give your applications a real‑world voice? Sign up for ElevenLabs today through this link: ElevenLabs. Grab your API key, spin up the script above, and start building immersive multi‑voice experiences right away. Happy coding!

Top comments (0)