DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Build a Real-Time Voice Translation App

Overview

Imagine a travel app that listens to a local conversation, instantly translates it, and speaks the result back in a natural‑sounding voice. With today’s cloud AI services you can pull this off in a few hundred lines of code. In this article we’ll walk through the core building blocks of a real‑time voice translation pipeline and show you how to glue them together with Python.

The stack we’ll use is:

Step Service / Library
Audio capture sounddevice (cross‑platform)
Speech‑to‑Text OpenAI Whisper (local model or API)
Translation Google Translate API (or the free googletrans wrapper)
Text‑to‑Speech ElevenLabs – natural‑sounding voice cloning (link below)

We’ll keep the example simple enough to run on a laptop, but the same pattern scales to serverless functions, mobile devices, or edge‑AI boxes.


Architecture

[Microphone] → Audio Stream → [Whisper] → Text → [Translate] → Target Text → [ElevenLabs TTS] → Playback
Enter fullscreen mode Exit fullscreen mode
  1. Capture short audio chunks (≈ 1 s) to keep latency low.
  2. Transcribe each chunk with Whisper.
  3. Translate the transcript into the target language.
  4. Synthesize the translated text using ElevenLabs’ high‑quality voice models.
  5. Play the audio back to the user.

Because each stage is stateless, you can run them in separate threads or even separate containers. Below we’ll focus on a single‑process prototype.


Capture Audio in Real Time

pyaudio works but sounddevice is lighter and pure‑Python. Install it first:

pip install sounddevice numpy
Enter fullscreen mode Exit fullscreen mode

The following snippet opens the default microphone, records 1‑second frames, and yields NumPy arrays:

import sounddevice as sd
import numpy as np

SAMPLE_RATE = 16_000  # Whisper expects 16 kHz mono
CHUNK_DURATION = 1.0  # seconds

def audio_generator():
    with sd.InputStream(samplerate=SAMPLE_RATE,
                        channels=1,
                        dtype='int16') as stream:
        while True:
            data, _ = stream.read(int(SAMPLE_RATE * CHUNK_DURATION))
            yield np.squeeze(data)
Enter fullscreen mode Exit fullscreen mode

The generator can be consumed by the next stage without blocking the UI.


Speech‑to‑Text with Whisper

You can either call OpenAI’s API or run the model locally. For a quick start, let’s use the whisper package:

pip install -U openai-whisper
Enter fullscreen mode Exit fullscreen mode
import whisper

# Load the tiny model – fast enough for real‑time on a modern laptop
model = whisper.load_model("tiny")

def transcribe(audio_chunk):
    # Whisper expects raw PCM bytes; convert from NumPy
    audio_bytes = audio_chunk.tobytes()
    result = model.transcribe(audio_bytes, language="en", fp16=False)
    return result["text"].strip()
Enter fullscreen mode Exit fullscreen mode

If you prefer the hosted API, replace the body with a simple requests.post to https://api.openai.com/v1/audio/transcriptions.


Translate the Text

Google’s free googletrans library works for prototypes. Install it:

pip install googletrans==4.0.0rc1
Enter fullscreen mode Exit fullscreen mode
from googletrans import Translator

translator = Translator()

def translate(text, target_lang="es"):
    # target_lang uses ISO 639‑1 codes, e.g., "es" for Spanish
    result = translator.translate(text, dest=target_lang)
    return result.text
Enter fullscreen mode Exit fullscreen mode

For production, swap in the official Cloud Translation API to get SLAs and quotas.


Text‑to‑Speech with ElevenLabs

ElevenLabs provides ultra‑natural voices and even lets you clone a custom voice in minutes. Sign up through the affiliate link to get free credits: ElevenLabs TTS API.

First, install the helper library:

pip install elevenlabs
Enter fullscreen mode Exit fullscreen mode

Then configure your API key (you’ll receive it after signing up):

import os
from elevenlabs import generate, play, set_api_key

set_api_key(os.getenv("ELEVENLABS_API_KEY"))
Enter fullscreen mode Exit fullscreen mode

A simple wrapper that turns translated text into audio:

def synthesize(text, voice="Rachel"):
    # voice can be any of ElevenLabs’ pre‑built voices or a custom ID
    audio = generate(
        text=text,
        voice=voice,
        model="eleven_multilingual_v2"  # multilingual model works for many languages
    )
    return audio
Enter fullscreen mode Exit fullscreen mode

You can also stream the result directly to the speaker using play(audio).


Putting It All Together

Below is a minimal end‑to‑end loop that captures audio, transcribes, translates to Spanish, synthesizes with ElevenLabs, and plays the result. Adjust TARGET_LANG and VOICE_NAME as needed.

import threading
import queue
import sounddevice as sd
import numpy as np
import whisper
from googletrans import Translator
from elevenlabs import generate, play, set_api_key
import os

# Config
SAMPLE_RATE = 16_000
CHUNK_DURATION = 1.0
TARGET_LANG = "es"   # Spanish
VOICE_NAME = "Rachel"

# Initialise services
model = whisper.load_model("tiny")
translator = Translator()
set_api_key(os.getenv("ELEVENLABS_API_KEY"))

def audio_producer(q):
    with sd.InputStream(samplerate=SAMPLE_RATE,
                        channels=1,
                        dtype='int16') as stream:
        while True:
            data, _ = stream.read(int(SAMPLE_RATE * CHUNK_DURATION))
            q.put(np.squeeze(data))

def processing_worker(q):
    while True:
        chunk = q.get()
        # 1️⃣ Speech‑to‑Text
        text = model.transcribe(chunk.tobytes(), language="en", fp16=False)["text"].strip()
        if not text:
            continue
        print(f"🗣️  Detected: {text}")

        # 2️⃣ Translate
        translated = translator.translate(text, dest=TARGET_LANG).text
        print(f"🔄 Translated ({TARGET_LANG}): {translated}")

        # 3️⃣ Synthesize
        audio = generate(text=translated, voice=VOICE_NAME, model="eleven_multilingual_v2")
        # 4️⃣ Play back
        play(audio)

if __name__ == "__main__":
    q = queue.Queue(maxsize=5)
    threading.Thread(target=audio_producer, args=(q,), daemon=True).start()
    threading.Thread(target=processing_worker, args=(q,), daemon=True).start()

    print("🚀 Real‑time voice translator is running. Press Ctrl+C to stop.")
    try:
        while True:
            pass
    except KeyboardInterrupt:
        print("\n👋 Bye!")
Enter fullscreen mode Exit fullscreen mode

What’s happening?

  • The audio_producer thread streams 1‑second PCM chunks into a bounded queue.
  • The processing_worker pulls each chunk, runs Whisper, translates, calls ElevenLabs, and finally plays the result.

Because each step is CPU‑light (tiny Whisper) or network‑bound (ElevenLabs), the overall latency stays under 2 seconds—acceptable for casual conversation.


Testing the Pipeline

  1. Set the API key
   export ELEVENLABS_API_KEY="your_api_key_here"
Enter fullscreen mode Exit fullscreen mode
  1. Run the script and speak a short English phrase (“How much does this cost?”).
  2. You should hear the Spanish equivalent (“¿Cuánto cuesta esto?”) spoken in a natural voice.

If you notice jitter, try increasing CHUNK_DURATION to 1.5 s or switch to a larger Whisper model (e.g., base). For production you’d also add error handling, back‑off retries, and maybe a small buffer to smooth playback.


Next Steps

  • Add language detection – Whisper can auto‑detect language; feed that into the translation step.
  • Support bidirectional conversation – Keep separate pipelines for inbound and outbound speech.
  • Deploy to the cloud – Wrap the logic in an HTTP endpoint (FastAPI) and let a mobile app stream audio via WebSockets.
  • Custom voice cloning – ElevenLabs lets you upload a few minutes of a speaker’s voice and generate a unique voice ID. Use that to give your app a brand‑specific voice.

Call to Action

Ready to give your users a truly natural listening experience? Jump straight into ElevenLabs’ high‑quality voice synthesis by signing up through the affiliate link: ElevenLabs TTS API – Get Started Here. With a few clicks you’ll have an API key, free credits, and access to the same voices that power the demo above. Happy coding!

Top comments (0)