Why a Voice‑Enabled Chatbot?
Chatbots have become the go‑to interface for customer support, personal assistants, and even hobby projects. Adding a voice layer turns a static text bot into a more natural, hands‑free experience. With the rise of powerful large‑language models (LLMs) and high‑quality text‑to‑speech (TTS) services, you can spin up a voice‑first assistant in a single afternoon.
In this guide we’ll stitch together OpenAI’s GPT‑4 for conversational intelligence and ElevenLabs for realistic speech synthesis. By the end you’ll have a Python script that listens to your microphone, sends the transcript to OpenAI, and speaks the response back using ElevenLabs’ voice cloning technology.
Tip: If you’re looking for a quick way to get lifelike speech, check out ElevenLabs here: https://try.elevenlabs.io/kr07zfuqn1bp
What You’ll Need
| Item | Reason |
|---|---|
| Python 3.9+ | Core language for the demo |
openai Python package |
Calls the GPT‑4 API |
pyaudio or sounddevice
|
Capture microphone audio |
requests |
Send HTTP requests to ElevenLabs |
| ElevenLabs API key | Access to their TTS endpoints |
| OpenAI API key | Talk to GPT‑4 |
You can install the required packages with:
pip install openai requests sounddevice numpy scipy
(If you prefer pyaudio, replace sounddevice with pyaudio.)
Setting Up the OpenAI Side
First, grab an API key from the OpenAI dashboard and store it securely, e.g. in an environment variable:
export OPENAI_API_KEY="sk-..."
A tiny helper function to query GPT‑4 looks like this:
import os
import openai
openai.api_key = os.getenv("OPENAI_API_KEY")
def ask_gpt(prompt: str) -> str:
response = openai.ChatCompletion.create(
model="gpt-4o-mini", # or "gpt-4" if you have access
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return response.choices[0].message["content"].strip()
Feel free to tweak temperature, max_tokens, or the model name to suit your use‑case.
Generating Speech with ElevenLabs
ElevenLabs provides a simple REST endpoint that accepts plain text and returns an MP3 (or WAV) audio stream. Sign up at the affiliate link to obtain an API key: https://try.elevenlabs.io/kr07zfuqn1bp
Here’s a minimal wrapper:
import os
import requests
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
VOICE_ID = "EXAVITQu4vr4xnSDxMaL" # default “Rachel” voice; replace with your cloned voice ID
def synthesize(text: str) -> bytes:
url = f"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1", # the high‑quality model
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
resp = requests.post(url, json=payload, headers=headers)
resp.raise_for_status()
return resp.content # raw audio bytes (MP3)
If you have a custom cloned voice, replace VOICE_ID with the ID you receive after uploading your voice sample.
Putting It All Together: A Simple Voice Loop
Below is a complete script that:
- Records a short utterance from your mic (≈ 3 seconds)
- Sends the audio to Whisper (via OpenAI) for transcription
- Passes the transcript to GPT‑4
- Sends GPT‑4’s reply to ElevenLabs for speech
- Plays the resulting audio back to you
import os, io, time, numpy as np, sounddevice as sd, scipy.io.wavfile as wav
import openai, requests
# ----- Config -----
openai.api_key = os.getenv("OPENAI_API_KEY")
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
VOICE_ID = "EXAVITQu4vr4xnSDxMaL" # change if you have a custom voice
# ----- Helper functions -----
def record(duration=3, fs=16000):
print("🎤 Listening…")
audio = sd.rec(int(duration * fs), samplerate=fs, channels=1, dtype='int16')
sd.wait()
return audio.squeeze()
def transcribe(audio_np):
# Convert numpy array to WAV bytes for Whisper
buf = io.BytesIO()
wav.write(buf, 16000, audio_np)
buf.seek(0)
transcript = openai.Audio.transcribe("whisper-1", buf)
return transcript["text"]
def ask_gpt(prompt):
resp = openai.ChatCompletion.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return resp.choices[0].message["content"].strip()
def synthesize(text):
url = f"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}"
headers = {"xi-api-key": ELEVEN_API_KEY, "Content-Type": "application/json"}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {"stability": 0.75, "similarity_boost": 0.85}
}
r = requests.post(url, json=payload, headers=headers)
r.raise_for_status()
return r.content
def play(audio_bytes):
# ElevenLabs returns MP3; convert to raw PCM for playback
import pygame
pygame.mixer.init()
sound = pygame.mixer.Sound(io.BytesIO(audio_bytes))
sound.play()
while pygame.mixer.get_busy():
time.sleep(0.1)
# ----- Main loop -----
if __name__ == "__main__":
while True:
try:
# 1️⃣ Capture voice
raw = record()
# 2️⃣ Transcribe
user_text = transcribe(raw)
print(f"🗣️ You said: {user_text}")
# 3️⃣ Get LLM reply
reply = ask_gpt(user_text)
print(f"🤖 Bot: {reply}")
# 4️⃣ Convert reply to speech
audio = synthesize(reply)
# 5️⃣ Play back
play(audio)
except KeyboardInterrupt:
print("\n👋 Bye!")
break
except Exception as e:
print(f"❗ Error: {e}")
What’s happening under the hood?
- Whisper (OpenAI’s speech‑to‑text model) handles the transcription step. It’s surprisingly accurate for short commands and works directly with the raw audio we captured.
- GPT‑4 supplies the conversational intelligence. You can swap the model or add system prompts to shape the bot’s personality.
- ElevenLabs renders the text into a natural‑sounding voice. Their “stability” and “similarity_boost” parameters let you fine‑tune how expressive vs. precise the speech sounds.
Going Further
Real‑time Streaming
If you need lower latency, consider streaming audio to Whisper via the audio.transcriptions endpoint and using ElevenLabs’ streaming TTS (available in their beta). The pattern stays the same—just replace the blocking requests.post with a websocket client.
Deploying as a Serverless Function
Wrap the ask_gpt and synthesize calls into an HTTP endpoint (e.g., FastAPI or AWS Lambda). Front‑end apps can then send text or audio payloads and receive a URL to an MP3 that can be streamed directly in the browser.
Voice Cloning
ElevenLabs shines when you upload a few seconds of your own voice and let the service generate a personalized voice ID. Use the same synthesize function; just swap VOICE_ID with the ID returned after the cloning process. The result feels like you’re talking to yourself!
Wrap‑Up
You now have a fully functional voice‑enabled chatbot built with OpenAI and ElevenLabs. The core ideas—record, transcribe, generate, synthesize—are reusable across many domains: virtual assistants, language learning tools, accessibility apps, and more.
If you enjoyed the demo and want to experiment with higher‑quality voices or custom clones, give ElevenLabs a spin: https://try.elevenlabs.io/kr07zfuqn1bp. Their API is developer‑friendly, fast, and the speech quality is truly impressive. Happy coding!
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.