Why AI Dubbing Matters
If you’ve ever tried to repurpose a tutorial, vlog, or interview for a different language or audience, you know how painful the manual dubbing process can be. You have to translate the script, record a voice‑over, line‑up the timing, and then re‑encode the video. With modern text‑to‑speech (TTS) and voice cloning services, you can automate most of that pipeline in a few lines of code.
In this article we’ll walk through building a lightweight AI dubbing tool that:
- Extracts the spoken transcript from a video (or uses an existing subtitle file).
- Sends the transcript to a TTS service that can clone a natural‑sounding voice.
- Merges the generated audio back into the original video while preserving sync.
The star of the show is ElevenLabs – a high‑quality TTS platform that offers voice cloning, SSML support, and a straightforward REST API. We’ll use the free tier for the demo, but the same code works for paid plans that give you longer audio limits and custom voice training.
Prerequisites
| What you need | Why |
|---|---|
| Python 3.9+ | For scripting the pipeline. |
ffmpeg installed and available in your PATH |
To extract audio, splice tracks, and re‑encode the final video. |
| An ElevenLabs API key | To call the TTS endpoint. |
| A source video (MP4, MOV, etc.) | The content you want to dub. |
You can install the Python dependencies with:
pip install requests tqdm python-dotenv
We’ll keep the code simple and avoid heavy frameworks – just requests for HTTP calls and tqdm for progress bars.
1. Grab the Transcript
If you already have an SRT or VTT file you can skip this step. Otherwise, let’s extract the audio and send it to a speech‑to‑text service. For the sake of brevity we’ll use the free Whisper API (or you can run whisper.cpp locally).
import subprocess
from pathlib import Path
def extract_audio(video_path: str, out_wav: str = "audio.wav"):
"""Extracts a mono 16‑kHz WAV file for best ASR results."""
cmd = [
"ffmpeg",
"-y", # overwrite output
"-i", video_path,
"-ac", "1", # mono
"-ar", "16000", # 16 kHz
out_wav
]
subprocess.run(cmd, check=True)
# Example usage
extract_audio("my_video.mp4")
Now you can feed audio.wav to Whisper (or any other ASR) and get a plain‑text transcript. The output should be a list of sentences with timestamps – we’ll need those timestamps later to keep the dubbing in sync.
2. Call ElevenLabs TTS
ElevenLabs offers two useful endpoints:
-
/v1/text-to-speech/{voice_id}– Generate speech from plain text. -
/v1/text-to-speech/{voice_id}/stream– Stream the audio directly (useful for large files).
Below is a minimal wrapper that sends a chunk of text and receives an MP3 file. We’ll split the transcript into sentences (or subtitle blocks) so we can later align each audio segment with its original timing.
import os
import requests
from tqdm import tqdm
from dotenv import load_dotenv
load_dotenv() # expects ELEVENLABS_API_KEY in .env
ELEVEN_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
def generate_speech(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> bytes:
"""
Sends `text` to ElevenLabs and returns raw MP3 bytes.
The default voice_id is ElevenLabs' public “Rachel” voice.
"""
url = f"{BASE_URL}/text-to-speech/{voice_id}"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
response = requests.post(url, json=payload, headers=headers, timeout=30)
response.raise_for_status()
return response.content
def batch_synthesize(sentences: list, out_dir: str = "tts_chunks"):
os.makedirs(out_dir, exist_ok=True)
for idx, s in enumerate(tqdm(sentences, desc="Synthesizing")):
audio_bytes = generate_speech(s["text"])
out_path = Path(out_dir) / f"{idx:04d}.mp3"
out_path.write_bytes(audio_bytes)
s["audio_path"] = str(out_path) # store for later stitching
Tip: If you have a custom cloned voice, replace voice_id with the ID you receive after uploading the voice sample to ElevenLabs. The same endpoint works for any voice you own.
3. Stitch Audio Back to Video
Now we have a directory of MP3 files, each representing a sentence. We need to place them on the timeline according to the timestamps we captured from the ASR step.
import json
import subprocess
def concat_audio_chunks(chunks: list, out_path: str = "dubbed.wav"):
"""
Concatenates MP3 chunks into a single WAV file respecting timestamps.
For simplicity we assume no gaps; you can add silence if needed.
"""
# Create a temporary file list for ffmpeg
list_file = "chunks.txt"
with open(list_file, "w") as f:
for ch in chunks:
f.write(f"file '{ch['audio_path']}'\n")
# Concatenate into WAV (16k mono) for easier video muxing
cmd = [
"ffmpeg", "-y", "-f", "concat", "-safe", "0",
"-i", list_file,
"-ac", "1", "-ar", "16000",
out_path
]
subprocess.run(cmd, check=True)
os.remove(list_file)
# Example usage after batch_synthesize()
# Assume `sentences` is a list of dicts: {"text": "...", "start": 1.23, "end": 2.45}
concat_audio_chunks(sentences, "dubbed.wav")
Finally, replace the original audio track with the newly generated dub:
def mux_video(original_video: str, new_audio: str, out_video: str = "video_dubbed.mp4"):
cmd = [
"ffmpeg", "-y",
"-i", original_video,
"-i", new_audio,
"-c:v", "copy", # keep original video codec
"-map", "0:v", # video from first input
"-map", "1:a", # audio from second input
"-shortest", # stop at shortest stream
out_video
]
subprocess.run(cmd, check=True)
# Run the final step
mux_video("my_video.mp4", "dubbed.wav", "my_video_dubbed.mp4")
That’s it! You now have a fully dubbed video with a synthetic voice that sounds surprisingly human.
4. Polishing the Experience
Adding SSML for Emphasis
ElevenLabs supports a subset of SSML tags (e.g., <break>, <prosody>). If you want a pause before a new paragraph or a slightly higher pitch for a question, just embed the tags in the text you send:
payload = {
"text": "<speak>Hello there!<break time='500ms'/>How are you today?</speak>",
"model_id": "eleven_monolingual_v1",
# rest of the payload unchanged
}
Handling Long Videos
The free tier caps each request at ~30 seconds of audio. For longer content, split the transcript into smaller blocks (as we did) and call the API in a loop. Respect the rate limits (≈ 5 req/s) to avoid 429 errors.
Custom Voice Cloning
If you need a brand‑specific voice, ElevenLabs lets you upload a few minutes of clean speech and creates a clone. The workflow stays the same; just replace the default voice_id with the ID of your custom voice. You can read more about the cloning process on the ElevenLabs website.
5. Full Minimal Script
Putting the pieces together, here’s a concise end‑to‑end script you can drop into a repo:
# ai_dub.py
import json, os, subprocess
from pathlib import Path
from tqdm import tqdm
import requests
from dotenv import load_dotenv
load_dotenv()
ELEVEN_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
def extract_audio(video):
subprocess.run([
"ffmpeg", "-y", "-i", video,
"-ac", "1", "-ar", "16000", "audio.wav"
], check=True)
def transcribe_whisper(audio_path):
# Placeholder – replace with your preferred ASR call
# Return list of {"text": "...", "start": 0.0, "end": 2.3}
raise NotImplementedError
def generate_speech(text, voice_id="EXAVITQu4vr4xnSDxMaL"):
url = f"{BASE_URL}/text-to-speech/{voice_id}"
headers = {"xi-api-key": ELEVEN_API_KEY, "Content-Type": "application/json"}
payload = {"text": text, "model_id": "eleven_monolingual_v1"}
r = requests.post(url, json=payload, headers=headers)
r.raise_for_status()
return r.content
def synthesize_chunks(sentences):
out_dir = Path("chunks")
out_dir.mkdir(exist_ok=True)
for i, s in enumerate(tqdm(sentences)):
audio = generate_speech(s["text"])
p = out_dir / f"{i:04d}.mp3"
p.write_bytes(audio)
s["audio_path"] = str(p)
def concat_chunks(sentences):
list_file = "list.txt"
with open(list_file, "w") as f:
for s in sentences:
f.write(f"file '{s['audio_path']}'\n")
subprocess.run([
"ffmpeg", "-y", "-f", "concat", "-safe", "0",
"-i", list_file, "-ac", "1", "-ar", "16000", "dubbed.wav"
], check=True)
os.remove(list_file)
def mux(video):
subprocess.run([
"ffmpeg", "-y", "-i", video, "-i", "dubbed.wav",
"-c:v", "copy", "-map", "0:v", "-map", "1:a",
"-shortest", "video_dubbed.mp4"
], check=True)
if __name__ == "__main__":
src_video = "my_video.mp4"
extract_audio(src_video)
# Replace with real transcription logic
# sentences = transcribe_whisper("audio.wav")
# For demo, we mock a single sentence:
sentences = [{"text": "Hello world!", "start": 0, "end": 2}]
synthesize_chunks(sentences)
concat_chunks(sentences)
mux(src_video)
Feel free to expand the script with proper error handling, parallel TTS calls, or integration with a subtitle editor.
Wrap‑Up
Building an AI dubbing pipeline is now a matter of stitching together a few well‑documented pieces:
- Extract the audio and get a timestamped transcript.
- Generate high‑quality speech with a service like ElevenLabs, optionally using a custom cloned voice.
-
Merge the new audio back into the video with
ffmpeg.
The result is a fully dubbed video ready for distribution, localization, or rapid prototyping of multilingual content.
Ready to give it a spin?
Grab your free ElevenLabs API key, plug the link into the script above, and start dubbing your own videos today. The more you experiment with SSML and voice settings, the closer the output will feel to a human narrator.
Try ElevenLabs now and bring your video content to life with AI‑powered voice! https://try.elevenlabs.io/kr07zfuqn1bp
Top comments (0)