Why Combine ElevenLabs and Twilio?
If you’ve ever wanted to turn a simple phone call into an interactive, AI‑powered experience, you’re in the right place. Twilio gives you the plumbing to make and receive voice calls, while ElevenLabs provides state‑of‑the‑art text‑to‑speech (TTS) and voice cloning. Put them together and you can build everything from personalized voicemail assistants to real‑time language translation over the phone.
In this post we’ll walk through a minimal but fully functional example:
- Receive an inbound call with Twilio
-
Gather the caller’s speech (using Twilio’s
<Gather>) - Send the transcript to OpenAI’s Whisper (or any STT you prefer)
- Generate a response with ChatGPT
- Render the response as natural‑sounding audio with ElevenLabs
- Play the audio back to the caller
By the end you’ll have a reusable Flask endpoint that you can deploy on any cloud provider.
Pro tip: If you need a high‑quality voice for your brand, ElevenLabs’ voice cloning lets you upload a few minutes of audio and get a custom voice that sounds like a real person. Check it out here: https://try.elevenlabs.io/kr07zfuqn1bp
Prerequisites
| What you need | Why |
|---|---|
| Twilio account | To get a phone number and expose a webhook |
| ElevenLabs API key | To call the TTS endpoint |
| OpenAI API key (optional) | For speech‑to‑text and conversational AI |
| Python 3.9+ | The example uses Flask |
| ngrok (or any public URL) | To expose your local server to Twilio |
Install the required Python packages:
pip install flask twilio requests openai
Step 1 – Set Up the Twilio Webhook
Create a new phone number in the Twilio console and point its Voice & Fax → A CALL COMES IN webhook to https://<your‑public‑url>/voice. When Twilio receives a call it will POST an XML document (Twiml) that tells it what to do next.
# app.py
from flask import Flask, request, Response
from twilio.twiml.voice_response import VoiceResponse, Gather
app = Flask(__name__)
@app.route("/voice", methods=["POST"])
def voice():
resp = VoiceResponse()
# Ask the caller to speak after the beep
gather = Gather(input="speech", action="/process", method="POST", timeout=5)
gather.say("Hi! Tell me what you need help with.")
resp.append(gather)
# If no speech was captured, fallback
resp.say("Sorry, I didn't catch that. Goodbye!")
resp.hangup()
return Response(str(resp), mimetype="application/xml")
Run the app locally and expose it with ngrok:
python app.py
ngrok http 5000
Take the HTTPS ngrok URL (e.g., https://abcd1234.ngrok.io) and paste it into the Twilio webhook field.
Step 2 – Capture Speech and Turn It Into Text
When the caller finishes speaking, Twilio sends a POST to /process with a SpeechResult field that already contains a transcription (thanks to Twilio’s built‑in speech recognition). If you want higher accuracy or multiple languages, you can pipe the raw audio to Whisper instead, but for this demo we’ll use Twilio’s result directly.
import os
import openai
import requests
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
ELEVEN_TTS_URL = "https://api.elevenlabs.io/v1/text-to-speech"
openai.api_key = OPENAI_API_KEY
@app.route("/process", methods=["POST"])
def process():
# Grab what Twilio recognized
user_text = request.form.get("SpeechResult", "")
if not user_text:
return fallback_response("I didn't hear anything.")
# Generate a ChatGPT response
reply = chatgpt_reply(user_text)
# Convert reply to audio with ElevenLabs
audio_url = eleven_tts(reply)
# Build TwiML to play the audio
resp = VoiceResponse()
resp.play(audio_url)
resp.hangup()
return Response(str(resp), mimetype="application/xml")
Helper: ChatGPT Reply
def chatgpt_reply(prompt: str) -> str:
completion = openai.ChatCompletion.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return completion.choices[0].message.content.strip()
Step 3 – Send the Text to ElevenLabs for TTS
ElevenLabs’ API accepts plain text and returns a short‑lived URL to an MP3 file. The endpoint we’ll use is /v1/text-to-speech/{voice_id}. For most developers the “default” voice (EXAVITQu4vr4xnSDxMaL) works great, but you can replace it with a cloned voice ID if you’ve uploaded your own samples.
def eleven_tts(text: str) -> str:
voice_id = "EXAVITQu4vr4xnSDxMaL" # default voice
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
response = requests.post(
f"{ELEVEN_TTS_URL}/{voice_id}",
json=payload,
headers=headers,
timeout=15,
)
response.raise_for_status()
# The API returns raw audio bytes; we upload to a temporary storage service
# For simplicity, we’ll use ElevenLabs’ own temporary URL feature:
return response.json()["audio_url"]
Note: If you prefer not to host the MP3 yourself, ElevenLabs can give you a temporary URL that Twilio can stream directly, as shown above.
Step 4 – Put It All Together
Your final app.py should look roughly like this:
import os
from flask import Flask, request, Response
from twilio.twiml.voice_response import VoiceResponse, Gather
import openai
import requests
app = Flask(__name__)
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
ELEVEN_TTS_URL = "https://api.elevenlabs.io/v1/text-to-speech"
openai.api_key = OPENAI_API_KEY
def chatgpt_reply(prompt: str) -> str:
completion = openai.ChatCompletion.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return completion.choices[0].message.content.strip()
def eleven_tts(text: str) -> str:
voice_id = "EXAVITQu4vr4xnSDxMaL"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {"stability": 0.75, "similarity_boost": 0.85}
}
resp = requests.post(f"{ELEVEN_TTS_URL}/{voice_id}", json=payload, headers=headers)
resp.raise_for_status()
return resp.json()["audio_url"]
def fallback_response(message: str):
resp = VoiceResponse()
resp.say(message)
resp.hangup()
return Response(str(resp), mimetype="application/xml")
@app.route("/voice", methods=["POST"])
def voice():
resp = VoiceResponse()
gather = Gather(input="speech", action="/process", method="POST", timeout=5)
gather.say("Hey there! What can I help you with today?")
resp.append(gather)
resp.say("Sorry, I didn't hear anything. Bye!")
resp.hangup()
return Response(str(resp), mimetype="application/xml")
@app.route("/process", methods=["POST"])
def process():
user_text = request.form.get("SpeechResult", "")
if not user_text:
return fallback_response("I didn't catch that.")
reply = chatgpt_reply(user_text)
audio_url = eleven_tts(reply)
resp = VoiceResponse()
resp.play(audio_url)
resp.hangup()
return Response(str(resp), mimetype="application/xml")
if __name__ == "__main__":
app.run(debug=True, port=5000)
Deploy the script to your favorite host (Heroku, Fly.io, Render, etc.), point the Twilio webhook at the live URL, and you’re ready to make calls that sound like a real person—powered by ElevenLabs.
Going Further
| Feature | How to add it |
|---|---|
| Voice cloning | Upload a few minutes of your own voice to ElevenLabs, grab the returned voice_id, and replace the default voice_id in eleven_tts. |
| Multi‑language support | Use Whisper or Azure Speech‑to‑Text for transcription, then set model_id to eleven_multilingual_v2 when calling ElevenLabs. |
| Persisted conversation | Store the conversation_id from OpenAI and pass it back on each turn to maintain context. |
| Call recording | Add <Record> in the TwiML to capture the whole interaction for compliance or analytics. |
| Interactive menus | Use multiple <Gather> blocks and DTMF (input="dtmf") to build IVR trees. |
Debugging Tips
- Twilio logs – The console’s Debugger tab shows every request and response. Look for “Invalid TwiML” errors if the call hangs.
- ElevenLabs rate limits – Free tier caps at ~10,000 characters per month. If you hit a 429, throttle your requests or upgrade.
- Audio latency – ElevenLabs generates the MP3 in ~1‑2 seconds. For ultra‑low latency, pre‑generate common prompts or cache them in a CDN.
Wrap‑Up
Building a voice‑first AI app used to require a deep dive into DSP libraries and self‑hosted TTS engines. With Twilio handling the telephony plumbing and ElevenLabs delivering crystal‑clear, human‑like speech, the barrier to entry is now a few lines of code.
Ready to give your callers a voice that sounds truly human? Grab an API key from ElevenLabs and start experimenting today:
👉 Try ElevenLabs now: https://try.elevenlabs.io/kr07zfuqn1bp
Happy coding, and may your next voice app sound as good as it feels!
Top comments (0)