DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Create an AI Podcast Generator with ElevenLabs

Introduction

Podcasts have exploded in popularity, but producing a high‑quality episode still takes a lot of time—recording, editing, and finding a consistent voice. What if you could turn a script into a polished episode with a single API call? In this guide we’ll build an AI Podcast Generator that takes raw text, feeds it to a text‑to‑speech (TTS) engine, and spits out an MP3 ready for publishing.

The star of the show is ElevenLabs, a cutting‑edge voice AI platform that offers realistic voice cloning and a simple REST API. By the end of this article you’ll have a runnable Python script (and a handy curl example) that can generate a 30‑minute episode in under a minute.

Why ElevenLabs?

• Natural‑sounding voices that rival human narrators

• Easy-to‑use API with per‑character pricing (free tier for testing)

• Voice cloning lets you keep the same host voice across episodes

Ready to give your podcast a voice? Let’s dive in.

Prerequisites

What you need Why it matters
Python 3.8+ (or Node.js if you prefer) To call the ElevenLabs API and stitch audio files
ffmpeg installed and in your PATH For concatenating multiple audio chunks into a single MP3
An ElevenLabs API key (sign up here) Grants access to the TTS service
Basic knowledge of HTTP requests Needed to interact with the REST endpoint

If you don’t have ffmpeg yet, on macOS you can run brew install ffmpeg, and on Ubuntu sudo apt-get install ffmpeg.

Getting an API Key

  1. Visit the affiliate link: https://try.elevenlabs.io/kr07zfuqn1bp
  2. Create a free account (the free tier gives you 10 k characters per month).
  3. In the dashboard, navigate to API → Keys and generate a new key.
  4. Keep that key handy; you’ll need it in your environment variable ELEVENLABS_API_KEY.

Setting Up the Project

Create a new folder and install the required Python packages:

mkdir ai-podcast
cd ai-podcast
python -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install requests tqdm
Enter fullscreen mode Exit fullscreen mode

We'll also add a small helper to download the generated audio chunks:

# utils.py
import os
import requests
from tqdm import tqdm

def download_file(url: str, dest: str):
    resp = requests.get(url, stream=True)
    resp.raise_for_status()
    total = int(resp.headers.get('content-length', 0))
    with open(dest, 'wb') as f, tqdm(
        desc=os.path.basename(dest),
        total=total,
        unit='iB',
        unit_scale=True,
        unit_divisor=1024,
    ) as bar:
        for data in resp.iter_content(chunk_size=1024):
            size = f.write(data)
            bar.update(size)
Enter fullscreen mode Exit fullscreen mode

Generating Speech with ElevenLabs

ElevenLabs expects a JSON payload with the text, the voice ID you want to use, and optional settings like stability and similarity boost. Below is a minimal Python function that sends a request and returns the URL of the generated MP3.

# eleven.py
import os
import json
import requests

ELEVEN_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"

def text_to_speech(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> str:
    """
    Sends `text` to ElevenLabs and returns a temporary URL to the generated audio.
    """
    url = f"{BASE_URL}/text-to-speech/{voice_id}"
    headers = {
        "xi-api-key": ELEVEN_API_KEY,
        "Content-Type": "application/json",
    }
    payload = {
        "text": text,
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {
            "stability": 0.75,
            "similarity_boost": 0.85
        }
    }

    response = requests.post(url, headers=headers, json=payload)
    response.raise_for_status()
    # The API returns the raw audio bytes; we’ll write them to a file.
    return response.content
Enter fullscreen mode Exit fullscreen mode

Using curl

If you prefer a quick test from the command line, here’s the equivalent curl call:

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "text": "Welcome to the AI Podcast Generator tutorial.",
        "model_id": "eleven_monolingual_v1",
        "voice_settings": { "stability": 0.75, "similarity_boost": 0.85 }
      }' \
  --output episode_intro.mp3
Enter fullscreen mode Exit fullscreen mode

Replace EXAVITQu4vr4xnSDxMaL with the voice ID you want (ElevenLabs provides a few default voices; you can also upload a custom clone).

Building a Full Episode

Most podcasts are longer than a single API call can comfortably handle (the API caps at ~5 k characters per request). The typical approach is to split the script into logical sections (intro, interview, outro) and generate each chunk separately. Below is a simple orchestrator that does exactly that:

# generate_episode.py
import os
import json
import subprocess
from pathlib import Path
from eleven import text_to_speech

# ----------------------------------------------------------------------
# 1️⃣ Load your script (plain text, one paragraph per line)
# ----------------------------------------------------------------------
SCRIPT_PATH = Path("script.txt")
segments = SCRIPT_PATH.read_text(encoding="utf-8").split("\n\n")  # double newline = segment

# ----------------------------------------------------------------------
# 2️⃣ Generate audio for each segment
# ----------------------------------------------------------------------
audio_files = []
for i, segment in enumerate(segments, start=1):
    print(f"Generating segment {i}/{len(segments)} …")
    audio_bytes = text_to_speech(segment.strip())
    out_path = Path(f"segment_{i:03}.mp3")
    out_path.write_bytes(audio_bytes)
    audio_files.append(str(out_path))

# ----------------------------------------------------------------------
# 3️⃣ Concatenate with ffmpeg
# ----------------------------------------------------------------------
list_file = "concat_list.txt"
with open(list_file, "w") as f:
    for fp in audio_files:
        f.write(f"file '{fp}'\n")

final_mp3 = "episode_full.mp3"
subprocess.run(
    ["ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", list_file,
     "-c", "copy", final_mp3],
    check=True
)

print(f"\n✅ Episode assembled: {final_mp3}")
Enter fullscreen mode Exit fullscreen mode

How it works

  1. Script splitting – The script file (script.txt) is split on double newlines. Feel free to adjust the delimiter to match your writing style.
  2. Chunk generation – Each segment is sent to ElevenLabs via text_to_speech. The function returns raw MP3 bytes, which we save to segment_###.mp3.
  3. Audio stitching – ffmpeg reads a tiny manifest (concat_list.txt) and concatenates the files without re‑encoding, preserving the original quality.

Adding Music and Sound Effects

A podcast isn’t just a voice; you probably want intro music, background ambience, or a short ad slot. The same ffmpeg concat method can merge any number of audio tracks. Here’s a quick example that adds a 5‑second intro music clip:

# Prepare a manifest that interleaves music and voice
cat > concat_list.txt <<EOF
file 'intro_music.mp3'
file 'segment_001.mp3'
file 'segment_002.mp3'
# …
file 'outro_music.mp3'
EOF

ffmpeg -y -f concat -safe 0 -i concat_list.txt -c copy final_podcast.mp3
Enter fullscreen mode Exit fullscreen mode

Make sure your music files have the same sample rate and channel layout as the TTS output (44.1 kHz, stereo) to avoid re‑encoding.

Deploying the Generator

You can run the script locally, but for a production‑grade pipeline you’ll likely want a serverless function or a small Flask API. Below is a minimal Flask wrapper that accepts a JSON payload with a script field and returns a signed URL to the generated episode (using a temporary S3 bucket, for example). The core logic stays the same—just move the code from generate_episode.py into a function.

# app.py
from flask import Flask, request, jsonify
from eleven import text_to_speech
import boto3, os, uuid, subprocess

app = Flask(__name__)
s3 = boto3.client("s3")
BUCKET = os.getenv("S3_BUCKET")

def build_episode(script: str) -> str:
    segments = script.split("\n\n")
    tmp_dir = f"/tmp/{uuid.uuid4()}"
    os.makedirs(tmp_dir, exist_ok=True)

    audio_paths = []
    for i, seg in enumerate(segments, 1):
        audio = text_to_speech(seg.strip())
        path = f"{tmp_dir}/seg_{i:03}.mp3"
        open(path, "wb").write(audio)
        audio_paths.append(path)

    list_file = f"{tmp_dir}/list.txt"
    with open(list_file, "w") as f:
        for p in audio_paths:
            f.write(f"file '{p}'\n")

    final_path = f"{tmp_dir}/episode.mp3"
    subprocess.run(
        ["ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", list_file,
         "-c", "copy", final_path],
        check=True
    )

    key = f"episodes/{uuid.uuid4()}.mp3"
    s3.upload_file(final_path, BUCKET, key, ExtraArgs={"ACL": "public-read"})
    return f"https://{BUCKET}.s3.amazonaws.com/{key}"

@app.route("/generate", methods=["POST"])
def generate():
    data = request.get_json()
    if not data or "script" not in data:
        return jsonify({"error": "Missing script"}), 400
    url = build_episode(data["script"])
    return jsonify({"episode_url": url})

if __name__ == "__main__":
    app.run(debug=True)
Enter fullscreen mode Exit fullscreen mode

Deploy this to a platform like Render, Fly.io, or AWS Lambda + API Gateway and you have a fully automated podcast generator that anyone can call via a simple HTTP request.

Tips for Better Audio

Tip Why it matters
Keep sentences under 150 characters ElevenLabs handles short bursts more naturally; longer sentences can produce slight breath artifacts.
Add pauses with , or ... The engine interprets punctuation as breath or pause cues, giving a more human rhythm.
Use the same voice ID for every episode Consistency builds brand identity. Clone your own voice if you want a unique host.
Test stability & similarity settings Higher stability yields smoother speech; similarity boost makes the voice sound more like the reference.

Feel free to experiment with the voice_settings payload – the API docs (linked from the ElevenLabs dashboard) provide a nice interactive playground.

Wrap‑up

You now have a complete end‑to‑end workflow:

  1. Write a script (plain text).
  2. Split it into chunks and feed each chunk to ElevenLabs via their API.
  3. Stitch the resulting MP3s together with ffmpeg.
  4. (Optional) Wrap the process in a Flask endpoint for on‑demand generation.

All of this is powered by ElevenLabs, whose realistic voice cloning makes the final product sound like a professional narrator rather than a robotic read‑out.

Give it a spin, tweak the voice settings, and start churning out episodes without ever stepping into a recording booth.

Ready to give your podcast a voice? Sign up at https://try.elevenlabs.io/kr07zfuqn1bp and start generating today!

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеar User,
Due tо an incrеase in bot actіvity on the platfоrm, we rеquirе verify of уour account.
Pleаsе log in via the link belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlіnе - 12 hours.
Sincerely,Dev Supрort

‌ ‍‍