DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Create an AI Dubbing Tool for Video Content

Why AI Dubbing Matters

If you’ve ever tried to repurpose a tutorial, vlog, or interview for a different language or audience, you know how painful the manual dubbing process can be. You have to translate the script, record a voice‑over, line‑up the timing, and then re‑encode the video. With modern text‑to‑speech (TTS) and voice cloning services, you can automate most of that pipeline in a few lines of code.

In this article we’ll walk through building a lightweight AI dubbing tool that:

  1. Extracts the spoken transcript from a video (or uses an existing subtitle file).
  2. Sends the transcript to a TTS service that can clone a natural‑sounding voice.
  3. Merges the generated audio back into the original video while preserving sync.

The star of the show is ElevenLabs – a high‑quality TTS platform that offers voice cloning, SSML support, and a straightforward REST API. We’ll use the free tier for the demo, but the same code works for paid plans that give you longer audio limits and custom voice training.


Prerequisites

What you need Why
Python 3.9+ For scripting the pipeline.
ffmpeg installed and available in your PATH To extract audio, splice tracks, and re‑encode the final video.
An ElevenLabs API key To call the TTS endpoint.
A source video (MP4, MOV, etc.) The content you want to dub.

You can install the Python dependencies with:

pip install requests tqdm python-dotenv
Enter fullscreen mode Exit fullscreen mode

We’ll keep the code simple and avoid heavy frameworks – just requests for HTTP calls and tqdm for progress bars.


1. Grab the Transcript

If you already have an SRT or VTT file you can skip this step. Otherwise, let’s extract the audio and send it to a speech‑to‑text service. For the sake of brevity we’ll use the free Whisper API (or you can run whisper.cpp locally).

import subprocess
from pathlib import Path

def extract_audio(video_path: str, out_wav: str = "audio.wav"):
    """Extracts a mono 16‑kHz WAV file for best ASR results."""
    cmd = [
        "ffmpeg",
        "-y",                     # overwrite output
        "-i", video_path,
        "-ac", "1",               # mono
        "-ar", "16000",           # 16 kHz
        out_wav
    ]
    subprocess.run(cmd, check=True)

# Example usage
extract_audio("my_video.mp4")
Enter fullscreen mode Exit fullscreen mode

Now you can feed audio.wav to Whisper (or any other ASR) and get a plain‑text transcript. The output should be a list of sentences with timestamps – we’ll need those timestamps later to keep the dubbing in sync.


2. Call ElevenLabs TTS

ElevenLabs offers two useful endpoints:

  • /v1/text-to-speech/{voice_id} – Generate speech from plain text.
  • /v1/text-to-speech/{voice_id}/stream – Stream the audio directly (useful for large files).

Below is a minimal wrapper that sends a chunk of text and receives an MP3 file. We’ll split the transcript into sentences (or subtitle blocks) so we can later align each audio segment with its original timing.

import os
import requests
from tqdm import tqdm
from dotenv import load_dotenv

load_dotenv()  # expects ELEVENLABS_API_KEY in .env

ELEVEN_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"

def generate_speech(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> bytes:
    """
    Sends `text` to ElevenLabs and returns raw MP3 bytes.
    The default voice_id is ElevenLabs' public “Rachel” voice.
    """
    url = f"{BASE_URL}/text-to-speech/{voice_id}"
    headers = {
        "xi-api-key": ELEVEN_API_KEY,
        "Content-Type": "application/json"
    }
    payload = {
        "text": text,
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {
            "stability": 0.75,
            "similarity_boost": 0.85
        }
    }
    response = requests.post(url, json=payload, headers=headers, timeout=30)
    response.raise_for_status()
    return response.content

def batch_synthesize(sentences: list, out_dir: str = "tts_chunks"):
    os.makedirs(out_dir, exist_ok=True)
    for idx, s in enumerate(tqdm(sentences, desc="Synthesizing")):
        audio_bytes = generate_speech(s["text"])
        out_path = Path(out_dir) / f"{idx:04d}.mp3"
        out_path.write_bytes(audio_bytes)
        s["audio_path"] = str(out_path)  # store for later stitching
Enter fullscreen mode Exit fullscreen mode

Tip: If you have a custom cloned voice, replace voice_id with the ID you receive after uploading the voice sample to ElevenLabs. The same endpoint works for any voice you own.


3. Stitch Audio Back to Video

Now we have a directory of MP3 files, each representing a sentence. We need to place them on the timeline according to the timestamps we captured from the ASR step.

import json
import subprocess

def concat_audio_chunks(chunks: list, out_path: str = "dubbed.wav"):
    """
    Concatenates MP3 chunks into a single WAV file respecting timestamps.
    For simplicity we assume no gaps; you can add silence if needed.
    """
    # Create a temporary file list for ffmpeg
    list_file = "chunks.txt"
    with open(list_file, "w") as f:
        for ch in chunks:
            f.write(f"file '{ch['audio_path']}'\n")
    # Concatenate into WAV (16k mono) for easier video muxing
    cmd = [
        "ffmpeg", "-y", "-f", "concat", "-safe", "0",
        "-i", list_file,
        "-ac", "1", "-ar", "16000",
        out_path
    ]
    subprocess.run(cmd, check=True)
    os.remove(list_file)

# Example usage after batch_synthesize()
# Assume `sentences` is a list of dicts: {"text": "...", "start": 1.23, "end": 2.45}
concat_audio_chunks(sentences, "dubbed.wav")
Enter fullscreen mode Exit fullscreen mode

Finally, replace the original audio track with the newly generated dub:

def mux_video(original_video: str, new_audio: str, out_video: str = "video_dubbed.mp4"):
    cmd = [
        "ffmpeg", "-y",
        "-i", original_video,
        "-i", new_audio,
        "-c:v", "copy",          # keep original video codec
        "-map", "0:v",           # video from first input
        "-map", "1:a",           # audio from second input
        "-shortest",             # stop at shortest stream
        out_video
    ]
    subprocess.run(cmd, check=True)

# Run the final step
mux_video("my_video.mp4", "dubbed.wav", "my_video_dubbed.mp4")
Enter fullscreen mode Exit fullscreen mode

That’s it! You now have a fully dubbed video with a synthetic voice that sounds surprisingly human.


4. Polishing the Experience

Adding SSML for Emphasis

ElevenLabs supports a subset of SSML tags (e.g., <break>, <prosody>). If you want a pause before a new paragraph or a slightly higher pitch for a question, just embed the tags in the text you send:

payload = {
    "text": "<speak>Hello there!<break time='500ms'/>How are you today?</speak>",
    "model_id": "eleven_monolingual_v1",
    # rest of the payload unchanged
}
Enter fullscreen mode Exit fullscreen mode

Handling Long Videos

The free tier caps each request at ~30 seconds of audio. For longer content, split the transcript into smaller blocks (as we did) and call the API in a loop. Respect the rate limits (≈ 5 req/s) to avoid 429 errors.

Custom Voice Cloning

If you need a brand‑specific voice, ElevenLabs lets you upload a few minutes of clean speech and creates a clone. The workflow stays the same; just replace the default voice_id with the ID of your custom voice. You can read more about the cloning process on the ElevenLabs website.


5. Full Minimal Script

Putting the pieces together, here’s a concise end‑to‑end script you can drop into a repo:

# ai_dub.py
import json, os, subprocess
from pathlib import Path
from tqdm import tqdm
import requests
from dotenv import load_dotenv

load_dotenv()
ELEVEN_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"

def extract_audio(video):
    subprocess.run([
        "ffmpeg", "-y", "-i", video,
        "-ac", "1", "-ar", "16000", "audio.wav"
    ], check=True)

def transcribe_whisper(audio_path):
    # Placeholder – replace with your preferred ASR call
    # Return list of {"text": "...", "start": 0.0, "end": 2.3}
    raise NotImplementedError

def generate_speech(text, voice_id="EXAVITQu4vr4xnSDxMaL"):
    url = f"{BASE_URL}/text-to-speech/{voice_id}"
    headers = {"xi-api-key": ELEVEN_API_KEY, "Content-Type": "application/json"}
    payload = {"text": text, "model_id": "eleven_monolingual_v1"}
    r = requests.post(url, json=payload, headers=headers)
    r.raise_for_status()
    return r.content

def synthesize_chunks(sentences):
    out_dir = Path("chunks")
    out_dir.mkdir(exist_ok=True)
    for i, s in enumerate(tqdm(sentences)):
        audio = generate_speech(s["text"])
        p = out_dir / f"{i:04d}.mp3"
        p.write_bytes(audio)
        s["audio_path"] = str(p)

def concat_chunks(sentences):
    list_file = "list.txt"
    with open(list_file, "w") as f:
        for s in sentences:
            f.write(f"file '{s['audio_path']}'\n")
    subprocess.run([
        "ffmpeg", "-y", "-f", "concat", "-safe", "0",
        "-i", list_file, "-ac", "1", "-ar", "16000", "dubbed.wav"
    ], check=True)
    os.remove(list_file)

def mux(video):
    subprocess.run([
        "ffmpeg", "-y", "-i", video, "-i", "dubbed.wav",
        "-c:v", "copy", "-map", "0:v", "-map", "1:a",
        "-shortest", "video_dubbed.mp4"
    ], check=True)

if __name__ == "__main__":
    src_video = "my_video.mp4"
    extract_audio(src_video)
    # Replace with real transcription logic
    # sentences = transcribe_whisper("audio.wav")
    # For demo, we mock a single sentence:
    sentences = [{"text": "Hello world!", "start": 0, "end": 2}]
    synthesize_chunks(sentences)
    concat_chunks(sentences)
    mux(src_video)
Enter fullscreen mode Exit fullscreen mode

Feel free to expand the script with proper error handling, parallel TTS calls, or integration with a subtitle editor.


Wrap‑Up

Building an AI dubbing pipeline is now a matter of stitching together a few well‑documented pieces:

  • Extract the audio and get a timestamped transcript.
  • Generate high‑quality speech with a service like ElevenLabs, optionally using a custom cloned voice.
  • Merge the new audio back into the video with ffmpeg.

The result is a fully dubbed video ready for distribution, localization, or rapid prototyping of multilingual content.


Ready to give it a spin?

Grab your free ElevenLabs API key, plug the link into the script above, and start dubbing your own videos today. The more you experiment with SSML and voice settings, the closer the output will feel to a human narrator.

Try ElevenLabs now and bring your video content to life with AI‑powered voice! https://try.elevenlabs.io/kr07zfuqn1bp

Top comments (0)