DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Build a Voicemail Generator with ElevenLabs API

Overview

When you’re building a contact‑center or a smart‑home system, the last thing you want is an empty voicemail box. A little voice‑AI can turn a static “no answer” message into a personalized, dynamic greeting that feels like a real person. In this post we’ll walk through how to create a Voicemail Generator using the ElevenLabs Text‑to‑Speech (TTS) API. We’ll cover everything from authentication to generating a voice‑cloned message, then bundle it into a simple Flask app that can be called via webhook or a REST endpoint.

The result? A lightweight service that lets you generate a voicemail audio file on the fly, using the same high‑quality voices you can clone with ElevenLabs. Let’s dive in.

Prerequisites

Item Description
Python 3.8+ For the example code
pip To install dependencies
ElevenLabs API key Sign up at https://try.elevenlabs.io/kr07zfuqn1bp
Basic knowledge of Flask We’ll expose a simple HTTP endpoint

Tip: If you’re new to ElevenLabs, the link above gives you a free trial with credit to test the API.

Setting up ElevenLabs

ElevenLabs offers a powerful, low‑latency TTS endpoint that supports voice cloning, speaker embeddings, and a large library of natural‑sounding voices. The API is REST‑based, so you can call it from any language.

Get your API key

  1. Visit https://try.elevenlabs.io/kr07zfuqn1bp.
  2. Create an account or log in.
  3. Navigate to the API Keys section and copy the key.

Store it in an environment variable for security:

export ELEVENLABS_API_KEY="sk_your_key_here"
Enter fullscreen mode Exit fullscreen mode

Quick API test

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/eleven_monolingual_v1" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text":"Hello, this is a test."}'
Enter fullscreen mode Exit fullscreen mode

You should receive an audio stream in the response body. Great! You’re ready to embed this into an app.

Generating Speech with Voice Cloning

ElevenLabs lets you clone a voice by providing a short audio clip. For a voicemail system, you might want to use a company‑wide voice (e.g., a receptionist) or a custom voice that matches your brand.

import requests
import os

API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"

def clone_voice(audio_file_path, voice_name):
    """Clone a new voice from an audio sample."""
    headers = {
        "xi-api-key": API_KEY,
        "Content-Type": "application/json",
    }
    data = {
        "voice_name": voice_name,
        "audio_url": None,  # We'll upload the file directly
    }
    # Upload the audio file first
    with open(audio_file_path, "rb") as f:
        files = {"file": f}
        upload_resp = requests.post(f"{BASE_URL}/audio/upload", files=files, headers={"xi-api-key": API_KEY})
    upload_resp.raise_for_status()
    audio_url = upload_resp.json()["url"]

    # Create the voice
    data["audio_url"] = audio_url
    resp = requests.post(f"{BASE_URL}/voices", json=data, headers=headers)
    resp.raise_for_status()
    return resp.json()["voice_id"]
Enter fullscreen mode Exit fullscreen mode

Remember: The cloned voice is stored in your ElevenLabs account and can be reused across calls. You’ll get a voice_id that you’ll pass to the TTS endpoint.

Building the Voicemail Generator

We’ll create a Flask service with a single endpoint: /voicemail. It accepts JSON containing a caller_name, a message, and an optional voice_id. The service will:

  1. Compose a greeting (e.g., “Hi, this is [caller_name]. I’m sorry I missed your call.”).
  2. Use ElevenLabs to synthesize the speech.
  3. Return the audio file as a bytes stream.
from flask import Flask, request, send_file, jsonify
import requests
import os
import io

app = Flask(__name__)

ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"

def synthesize_text(text, voice_id):
    headers = {
        "xi-api-key": ELEVENLABS_API_KEY,
        "Content-Type": "application/json",
    }
    payload = {
        "text": text,
        "voice_id": voice_id,
        "model_id": "eleven_monolingual_v1",
    }
    resp = requests.post(f"{BASE_URL}/text-to-speech/{voice_id}", json=payload, headers=headers, stream=True)
    resp.raise_for_status()
    return resp.content

@app.route("/voicemail", methods=["POST"])
def voicemail():
    data = request.json
    caller = data.get("caller_name", "Someone")
    message = data.get("message", "I couldn't answer your call.")
    voice_id = data.get("voice_id")

    if not voice_id:
        # Fallback to a default voice
        voice_id = "EXkTlv5x2jJ5fK8V2i9F"  # Replace with your own default voice ID

    full_text = f"Hi, this is {caller}. {message}"
    audio_bytes = synthesize_text(full_text, voice_id)

    return send_file(
        io.BytesIO(audio_bytes),
        mimetype="audio/mpeg",
        as_attachment=True,
        download_name="voicemail.mp3",
    )

if __name__ == "__main__":
    app.run(debug=True)
Enter fullscreen mode Exit fullscreen mode

How it works

  • Endpoint: POST /voicemail Request body:
  {
    "caller_name": "Alice",
    "message": "Sorry I missed your call, please leave a message after the tone.",
    "voice_id": "EXkTlv5x2jJ5fK8V2i9F"
  }
Enter fullscreen mode Exit fullscreen mode
  • Response: An MP3 file named voicemail.mp3.

The synthesize_text helper streams the audio directly from ElevenLabs, so you’re not holding large buffers in memory.

Putting It All Together

Now that we have the core logic, let’s test the service locally.

# Start the server
python app.py
Enter fullscreen mode Exit fullscreen mode

In another terminal, call the endpoint:

curl -X POST "http://localhost:5000/voicemail" \
  -H "Content-Type: application/json" \
  -d '{"caller_name":"Bob","message":"Please leave a message after the beep."}' \
  -o voicemail.mp3
Enter fullscreen mode Exit fullscreen mode

Open voicemail.mp3 with your favorite player – you should hear a natural‑sounding greeting. If you want to use a cloned voice, pass the voice_id you obtained earlier.

Deployment Tips

  • Containerization: Wrap the Flask app in a Dockerfile. ElevenLabs API calls are stateless, so you can scale horizontally.
  • Environment variables: Keep your API key secret by using secrets management in your cloud provider (e.g., AWS Secrets Manager, GCP Secret Manager).
  • Caching: If you frequently use the same message, cache the audio on the server or in a CDN to reduce API usage and latency.

Advanced Ideas

  1. Dynamic Voice Selection: Load a pool of voice IDs and pick one based on caller region or time of day.
  2. Speech Synthesis Markers: Use ElevenLabs’ SSML support to add pauses or emphasis.
  3. Integration with Twilio: Hook the /voicemail endpoint into a Twilio webhook so that when a call is missed, Twilio automatically plays the generated audio.

Conclusion

ElevenLabs’ TTS API gives developers the ability to create high‑quality, personalized voicemails with minimal effort. By cloning a voice and exposing a simple REST endpoint, you can turn any missed call into a brand‑consistent, engaging experience.

If you’re ready to give your voicemail system a voice upgrade, grab a free trial and start experimenting today. Sign up here: https://try.elevenlabs.io/kr07zfuqn1bp and let ElevenLabs bring your voicemails to life!

Top comments (0)