DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Build a Text-to-Speech App with ElevenLabs and Python

Why ElevenLabs for Text‑to‑Speech?

If you’ve ever wanted to turn a paragraph of markdown into a natural‑sounding voice, you’ve probably tried the free tier of a cloud TTS service and ended up with robotic‑sounding output. ElevenLabs (https://try.elevenlabs.io/kr07zfuqn1bp) has raised the bar with deep‑learning models that capture nuance, intonation, and even speaker style. The API is straightforward, the latency is low, and the pricing is generous enough for hobby projects. In this post we’ll walk through building a tiny Python app that sends text to ElevenLabs, streams back an MP3, and plays it back locally.


Prerequisites

What you need Why
Python 3.8+ The language we’ll use for the demo
requests library Simple HTTP client
pydub (optional) Convert the raw MP3 into a playable WAV
An ElevenLabs API key Authenticate your requests – you can grab one from the dashboard after signing up via the affiliate link above

You can install the Python dependencies with:

pip install requests pydub
# On macOS/Linux you may also need ffmpeg for pydub:
brew install ffmpeg   # macOS
sudo apt-get install ffmpeg  # Ubuntu/Debian
Enter fullscreen mode Exit fullscreen mode

Getting Your API Key

  1. Sign up at the ElevenLabs site using this link: https://try.elevenlabs.io/kr07zfuqn1bp.
  2. Once you’re logged in, head to API → Keys and generate a new key.
  3. Keep the key handy; we’ll reference it as an environment variable called ELEVENLABS_API_KEY.
export ELEVENLABS_API_KEY="your_secret_key_here"
Enter fullscreen mode Exit fullscreen mode

The Core Request – A Minimal Python Function

ElevenLabs exposes a /v1/text-to-speech/{voice_id} endpoint. The simplest call looks like this:

import os
import requests

API_KEY = os.getenv("ELEVENLABS_API_KEY")
VOICE_ID = "EXAVITQu4vr4xnSDxMaL"   # default "Rachel" voice; pick any from the UI

def synthesize(text: str, output_path: str = "output.mp3"):
    url = f"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}"
    headers = {
        "xi-api-key": API_KEY,
        "Content-Type": "application/json"
    }
    payload = {
        "text": text,
        "model_id": "eleven_monolingual_v1",   # the latest high‑quality model
        "voice_settings": {
            "stability": 0.75,
            "similarity_boost": 0.85
        }
    }

    response = requests.post(url, json=payload, headers=headers, stream=True)
    response.raise_for_status()

    # Write the streamed MP3 directly to disk
    with open(output_path, "wb") as f:
        for chunk in response.iter_content(chunk_size=8192):
            f.write(chunk)

    print(f"✅ Saved audio to {output_path}")
Enter fullscreen mode Exit fullscreen mode

How It Works

  • stream=True – ElevenLabs streams the audio back in chunks, which avoids loading the whole file into memory.
  • voice_settings – Tweak stability (how steady the voice stays) and similarity_boost (how close it sounds to the reference voice). Play around to get the vibe you want.
  • model_id – The eleven_monolingual_v1 model is optimized for English. For multilingual projects, switch to eleven_multilingual_v1.

Quick Test from the REPL

if __name__ == "__main__":
    sample = """\
    Hey there! This is a quick demo of ElevenLabs' Text‑to‑Speech API.
    Notice how the pauses feel natural, and the intonation matches the punctuation.
    """
    synthesize(sample, "demo.mp3")
Enter fullscreen mode Exit fullscreen mode

Run the script, then play demo.mp3 with your favorite media player. You should hear a smooth, human‑like narration.


Curl Alternative – Handy for Debugging

Sometimes you just want to poke the endpoint from the terminal. Here’s a one‑liner that does the same thing:

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "text": "Hello from ElevenLabs! This is a curl test.",
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {"stability":0.75,"similarity_boost":0.85}
      }' \
  --output hello.mp3
Enter fullscreen mode Exit fullscreen mode

If you see a nicely sized MP3, the API is reachable and your key is valid.


Adding Voice Cloning (Optional)

One of the coolest features of ElevenLabs is voice cloning – you can upload a few seconds of a speaker’s audio and generate a custom voice ID. The flow is:

  1. Upload a sample (/v1/voices/add).
  2. Wait for training (usually a minute).
  3. Use the returned voice_id in the TTS request.

Below is a condensed version that assumes you already have a WAV file called my_voice.wav.

def upload_voice(sample_path: str) -> str:
    url = "https://api.elevenlabs.io/v1/voices/add"
    headers = {"xi-api-key": API_KEY}
    files = {"sample_file": open(sample_path, "rb")}
    data = {"name": "MyCustomVoice"}

    resp = requests.post(url, headers=headers, files=files, data=data)
    resp.raise_for_status()
    voice_id = resp.json()["voice_id"]
    print(f"✅ New voice created: {voice_id}")
    return voice_id
Enter fullscreen mode Exit fullscreen mode

After you obtain voice_id, replace the constant VOICE_ID in the synthesize function and you’ll hear your voice speaking the supplied text. This is perfect for creating personalized audiobooks, in‑game NPC dialogue, or even a brand‑specific assistant.


Packaging It as a Tiny Flask Service

If you want to expose the TTS capability over HTTP (e.g., for a web front‑end), a few lines of Flask do the trick:

from flask import Flask, request, send_file, jsonify
app = Flask(__name__)

@app.route("/tts", methods=["POST"])
def tts_endpoint():
    data = request.get_json()
    text = data.get("text")
    if not text:
        return jsonify({"error": "Missing 'text' field"}), 400

    output_file = "temp.mp3"
    synthesize(text, output_file)
    return send_file(output_file, mimetype="audio/mpeg")

if __name__ == "__main__":
    app.run(port=5000, debug=True)
Enter fullscreen mode Exit fullscreen mode

Now a simple POST request like:

curl -X POST http://localhost:5000/tts \
  -H "Content-Type: application/json" \
  -d '{"text":"Hello from my Flask TTS service!"}' \
  --output result.mp3
Enter fullscreen mode Exit fullscreen mode

gives you an MP3 you can stream directly to a browser or mobile app.


Tips for Production‑Ready Usage

Tip Reason
Cache results Many applications reuse the same sentences (e.g., UI prompts). Store the MP3 on disk or S3 to avoid unnecessary API calls.
Rate‑limit handling ElevenLabs enforces per‑minute quotas. Wrap the request in a retry loop with exponential back‑off.
Secure the API key Never commit the key to source control. Use environment variables, secret managers, or Vault.
Monitor latency Record the time from request to first byte. If you notice spikes, consider a local queue (RabbitMQ, Redis) to smooth traffic.

Wrap‑Up

You now have a fully functional pipeline:

  1. Grab an API key from ElevenLabs (https://try.elevenlabs.io/kr07zfuqn1bp).
  2. Send text to the /v1/text-to-speech endpoint with a few lines of Python.
  3. Optionally clone a voice to make the output uniquely yours.
  4. Expose it via Flask, FastAPI, or any framework you prefer.

ElevenLabs makes the heavy lifting of voice synthesis feel like a single HTTP call, letting you focus on the product logic instead of acoustic modeling.


Ready to give it a spin?

Head over to ElevenLabs, grab your free API key, and start building the next generation of voice‑enabled apps. Happy coding! 🚀


Top comments (0)