Overview
Imagine a travel app that listens to a local conversation, instantly translates it, and speaks the result back in a natural‑sounding voice. With today’s cloud AI services you can pull this off in a few hundred lines of code. In this article we’ll walk through the core building blocks of a real‑time voice translation pipeline and show you how to glue them together with Python.
The stack we’ll use is:
| Step | Service / Library |
|---|---|
| Audio capture |
sounddevice (cross‑platform) |
| Speech‑to‑Text | OpenAI Whisper (local model or API) |
| Translation | Google Translate API (or the free googletrans wrapper) |
| Text‑to‑Speech | ElevenLabs – natural‑sounding voice cloning (link below) |
We’ll keep the example simple enough to run on a laptop, but the same pattern scales to serverless functions, mobile devices, or edge‑AI boxes.
Architecture
[Microphone] → Audio Stream → [Whisper] → Text → [Translate] → Target Text → [ElevenLabs TTS] → Playback
- Capture short audio chunks (≈ 1 s) to keep latency low.
- Transcribe each chunk with Whisper.
- Translate the transcript into the target language.
- Synthesize the translated text using ElevenLabs’ high‑quality voice models.
- Play the audio back to the user.
Because each stage is stateless, you can run them in separate threads or even separate containers. Below we’ll focus on a single‑process prototype.
Capture Audio in Real Time
pyaudio works but sounddevice is lighter and pure‑Python. Install it first:
pip install sounddevice numpy
The following snippet opens the default microphone, records 1‑second frames, and yields NumPy arrays:
import sounddevice as sd
import numpy as np
SAMPLE_RATE = 16_000 # Whisper expects 16 kHz mono
CHUNK_DURATION = 1.0 # seconds
def audio_generator():
with sd.InputStream(samplerate=SAMPLE_RATE,
channels=1,
dtype='int16') as stream:
while True:
data, _ = stream.read(int(SAMPLE_RATE * CHUNK_DURATION))
yield np.squeeze(data)
The generator can be consumed by the next stage without blocking the UI.
Speech‑to‑Text with Whisper
You can either call OpenAI’s API or run the model locally. For a quick start, let’s use the whisper package:
pip install -U openai-whisper
import whisper
# Load the tiny model – fast enough for real‑time on a modern laptop
model = whisper.load_model("tiny")
def transcribe(audio_chunk):
# Whisper expects raw PCM bytes; convert from NumPy
audio_bytes = audio_chunk.tobytes()
result = model.transcribe(audio_bytes, language="en", fp16=False)
return result["text"].strip()
If you prefer the hosted API, replace the body with a simple requests.post to https://api.openai.com/v1/audio/transcriptions.
Translate the Text
Google’s free googletrans library works for prototypes. Install it:
pip install googletrans==4.0.0rc1
from googletrans import Translator
translator = Translator()
def translate(text, target_lang="es"):
# target_lang uses ISO 639‑1 codes, e.g., "es" for Spanish
result = translator.translate(text, dest=target_lang)
return result.text
For production, swap in the official Cloud Translation API to get SLAs and quotas.
Text‑to‑Speech with ElevenLabs
ElevenLabs provides ultra‑natural voices and even lets you clone a custom voice in minutes. Sign up through the affiliate link to get free credits: ElevenLabs TTS API.
First, install the helper library:
pip install elevenlabs
Then configure your API key (you’ll receive it after signing up):
import os
from elevenlabs import generate, play, set_api_key
set_api_key(os.getenv("ELEVENLABS_API_KEY"))
A simple wrapper that turns translated text into audio:
def synthesize(text, voice="Rachel"):
# voice can be any of ElevenLabs’ pre‑built voices or a custom ID
audio = generate(
text=text,
voice=voice,
model="eleven_multilingual_v2" # multilingual model works for many languages
)
return audio
You can also stream the result directly to the speaker using play(audio).
Putting It All Together
Below is a minimal end‑to‑end loop that captures audio, transcribes, translates to Spanish, synthesizes with ElevenLabs, and plays the result. Adjust TARGET_LANG and VOICE_NAME as needed.
import threading
import queue
import sounddevice as sd
import numpy as np
import whisper
from googletrans import Translator
from elevenlabs import generate, play, set_api_key
import os
# Config
SAMPLE_RATE = 16_000
CHUNK_DURATION = 1.0
TARGET_LANG = "es" # Spanish
VOICE_NAME = "Rachel"
# Initialise services
model = whisper.load_model("tiny")
translator = Translator()
set_api_key(os.getenv("ELEVENLABS_API_KEY"))
def audio_producer(q):
with sd.InputStream(samplerate=SAMPLE_RATE,
channels=1,
dtype='int16') as stream:
while True:
data, _ = stream.read(int(SAMPLE_RATE * CHUNK_DURATION))
q.put(np.squeeze(data))
def processing_worker(q):
while True:
chunk = q.get()
# 1️⃣ Speech‑to‑Text
text = model.transcribe(chunk.tobytes(), language="en", fp16=False)["text"].strip()
if not text:
continue
print(f"🗣️ Detected: {text}")
# 2️⃣ Translate
translated = translator.translate(text, dest=TARGET_LANG).text
print(f"🔄 Translated ({TARGET_LANG}): {translated}")
# 3️⃣ Synthesize
audio = generate(text=translated, voice=VOICE_NAME, model="eleven_multilingual_v2")
# 4️⃣ Play back
play(audio)
if __name__ == "__main__":
q = queue.Queue(maxsize=5)
threading.Thread(target=audio_producer, args=(q,), daemon=True).start()
threading.Thread(target=processing_worker, args=(q,), daemon=True).start()
print("🚀 Real‑time voice translator is running. Press Ctrl+C to stop.")
try:
while True:
pass
except KeyboardInterrupt:
print("\n👋 Bye!")
What’s happening?
- The
audio_producerthread streams 1‑second PCM chunks into a bounded queue. - The
processing_workerpulls each chunk, runs Whisper, translates, calls ElevenLabs, and finally plays the result.
Because each step is CPU‑light (tiny Whisper) or network‑bound (ElevenLabs), the overall latency stays under 2 seconds—acceptable for casual conversation.
Testing the Pipeline
- Set the API key
export ELEVENLABS_API_KEY="your_api_key_here"
- Run the script and speak a short English phrase (“How much does this cost?”).
- You should hear the Spanish equivalent (“¿Cuánto cuesta esto?”) spoken in a natural voice.
If you notice jitter, try increasing CHUNK_DURATION to 1.5 s or switch to a larger Whisper model (e.g., base). For production you’d also add error handling, back‑off retries, and maybe a small buffer to smooth playback.
Next Steps
- Add language detection – Whisper can auto‑detect language; feed that into the translation step.
- Support bidirectional conversation – Keep separate pipelines for inbound and outbound speech.
- Deploy to the cloud – Wrap the logic in an HTTP endpoint (FastAPI) and let a mobile app stream audio via WebSockets.
- Custom voice cloning – ElevenLabs lets you upload a few minutes of a speaker’s voice and generate a unique voice ID. Use that to give your app a brand‑specific voice.
Call to Action
Ready to give your users a truly natural listening experience? Jump straight into ElevenLabs’ high‑quality voice synthesis by signing up through the affiliate link: ElevenLabs TTS API – Get Started Here. With a few clicks you’ll have an API key, free credits, and access to the same voices that power the demo above. Happy coding!
Top comments (0)