Overview
When you’re building a contact‑center or a smart‑home system, the last thing you want is an empty voicemail box. A little voice‑AI can turn a static “no answer” message into a personalized, dynamic greeting that feels like a real person. In this post we’ll walk through how to create a Voicemail Generator using the ElevenLabs Text‑to‑Speech (TTS) API. We’ll cover everything from authentication to generating a voice‑cloned message, then bundle it into a simple Flask app that can be called via webhook or a REST endpoint.
The result? A lightweight service that lets you generate a voicemail audio file on the fly, using the same high‑quality voices you can clone with ElevenLabs. Let’s dive in.
Prerequisites
| Item | Description |
|---|---|
| Python 3.8+ | For the example code |
pip |
To install dependencies |
| ElevenLabs API key | Sign up at https://try.elevenlabs.io/kr07zfuqn1bp |
| Basic knowledge of Flask | We’ll expose a simple HTTP endpoint |
Tip: If you’re new to ElevenLabs, the link above gives you a free trial with credit to test the API.
Setting up ElevenLabs
ElevenLabs offers a powerful, low‑latency TTS endpoint that supports voice cloning, speaker embeddings, and a large library of natural‑sounding voices. The API is REST‑based, so you can call it from any language.
Get your API key
- Visit https://try.elevenlabs.io/kr07zfuqn1bp.
- Create an account or log in.
- Navigate to the API Keys section and copy the key.
Store it in an environment variable for security:
export ELEVENLABS_API_KEY="sk_your_key_here"
Quick API test
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/eleven_monolingual_v1" \
-H "xi-api-key: $ELEVENLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Hello, this is a test."}'
You should receive an audio stream in the response body. Great! You’re ready to embed this into an app.
Generating Speech with Voice Cloning
ElevenLabs lets you clone a voice by providing a short audio clip. For a voicemail system, you might want to use a company‑wide voice (e.g., a receptionist) or a custom voice that matches your brand.
import requests
import os
API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
def clone_voice(audio_file_path, voice_name):
"""Clone a new voice from an audio sample."""
headers = {
"xi-api-key": API_KEY,
"Content-Type": "application/json",
}
data = {
"voice_name": voice_name,
"audio_url": None, # We'll upload the file directly
}
# Upload the audio file first
with open(audio_file_path, "rb") as f:
files = {"file": f}
upload_resp = requests.post(f"{BASE_URL}/audio/upload", files=files, headers={"xi-api-key": API_KEY})
upload_resp.raise_for_status()
audio_url = upload_resp.json()["url"]
# Create the voice
data["audio_url"] = audio_url
resp = requests.post(f"{BASE_URL}/voices", json=data, headers=headers)
resp.raise_for_status()
return resp.json()["voice_id"]
Remember: The cloned voice is stored in your ElevenLabs account and can be reused across calls. You’ll get a
voice_idthat you’ll pass to the TTS endpoint.
Building the Voicemail Generator
We’ll create a Flask service with a single endpoint: /voicemail. It accepts JSON containing a caller_name, a message, and an optional voice_id. The service will:
- Compose a greeting (e.g., “Hi, this is [caller_name]. I’m sorry I missed your call.”).
- Use ElevenLabs to synthesize the speech.
- Return the audio file as a
bytesstream.
from flask import Flask, request, send_file, jsonify
import requests
import os
import io
app = Flask(__name__)
ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
def synthesize_text(text, voice_id):
headers = {
"xi-api-key": ELEVENLABS_API_KEY,
"Content-Type": "application/json",
}
payload = {
"text": text,
"voice_id": voice_id,
"model_id": "eleven_monolingual_v1",
}
resp = requests.post(f"{BASE_URL}/text-to-speech/{voice_id}", json=payload, headers=headers, stream=True)
resp.raise_for_status()
return resp.content
@app.route("/voicemail", methods=["POST"])
def voicemail():
data = request.json
caller = data.get("caller_name", "Someone")
message = data.get("message", "I couldn't answer your call.")
voice_id = data.get("voice_id")
if not voice_id:
# Fallback to a default voice
voice_id = "EXkTlv5x2jJ5fK8V2i9F" # Replace with your own default voice ID
full_text = f"Hi, this is {caller}. {message}"
audio_bytes = synthesize_text(full_text, voice_id)
return send_file(
io.BytesIO(audio_bytes),
mimetype="audio/mpeg",
as_attachment=True,
download_name="voicemail.mp3",
)
if __name__ == "__main__":
app.run(debug=True)
How it works
-
Endpoint:
POST /voicemailRequest body:
{
"caller_name": "Alice",
"message": "Sorry I missed your call, please leave a message after the tone.",
"voice_id": "EXkTlv5x2jJ5fK8V2i9F"
}
-
Response: An MP3 file named
voicemail.mp3.
The synthesize_text helper streams the audio directly from ElevenLabs, so you’re not holding large buffers in memory.
Putting It All Together
Now that we have the core logic, let’s test the service locally.
# Start the server
python app.py
In another terminal, call the endpoint:
curl -X POST "http://localhost:5000/voicemail" \
-H "Content-Type: application/json" \
-d '{"caller_name":"Bob","message":"Please leave a message after the beep."}' \
-o voicemail.mp3
Open voicemail.mp3 with your favorite player – you should hear a natural‑sounding greeting. If you want to use a cloned voice, pass the voice_id you obtained earlier.
Deployment Tips
- Containerization: Wrap the Flask app in a Dockerfile. ElevenLabs API calls are stateless, so you can scale horizontally.
- Environment variables: Keep your API key secret by using secrets management in your cloud provider (e.g., AWS Secrets Manager, GCP Secret Manager).
- Caching: If you frequently use the same message, cache the audio on the server or in a CDN to reduce API usage and latency.
Advanced Ideas
- Dynamic Voice Selection: Load a pool of voice IDs and pick one based on caller region or time of day.
- Speech Synthesis Markers: Use ElevenLabs’ SSML support to add pauses or emphasis.
-
Integration with Twilio: Hook the
/voicemailendpoint into a Twilio webhook so that when a call is missed, Twilio automatically plays the generated audio.
Conclusion
ElevenLabs’ TTS API gives developers the ability to create high‑quality, personalized voicemails with minimal effort. By cloning a voice and exposing a simple REST endpoint, you can turn any missed call into a brand‑consistent, engaging experience.
If you’re ready to give your voicemail system a voice upgrade, grab a free trial and start experimenting today. Sign up here: https://try.elevenlabs.io/kr07zfuqn1bp and let ElevenLabs bring your voicemails to life!
Top comments (0)