DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Build an Audiobook Generator Using ElevenLabs API

Introduction

Ever wanted to turn a novel, a blog series, or a collection of technical docs into a polished audiobook—without spending hours in a recording studio? With modern voice‑AI services you can generate natural‑sounding narration programmatically, and ElevenLabs makes it surprisingly easy. In this walkthrough we’ll build a simple “Audiobook Generator” that pulls plain‑text chapters, sends them to ElevenLabs’ Text‑to‑Speech (TTS) API, and stitches the resulting MP3 files together into a single audiobook file.

You’ll walk away with a reusable Python script (and a quick curl example) that you can plug into any content pipeline, whether you’re a solo indie author, a SaaS that serves audio versions of articles, or just a hobbyist who loves to experiment with voice cloning.


Prerequisites

What you need Why
Python 3.8+ (or Node if you prefer) To drive the API calls and file handling
An ElevenLabs account (free tier works for short clips) Provides the TTS endpoint and voice models
ffmpeg installed and on your PATH Concatenates MP3 segments into a single file
Basic familiarity with REST APIs Helps you tweak request parameters

If you don’t have ffmpeg yet, install it via your package manager:

# macOS
brew install ffmpeg

# Ubuntu/Debian
sudo apt-get install ffmpeg
Enter fullscreen mode Exit fullscreen mode

Getting an API key

  1. Sign up at the ElevenLabs platform using this affiliate link: https://try.elevenlabs.io/kr07zfuqn1bp.
  2. Once logged in, navigate to API → Keys and generate a new key.
  3. Copy the key; you’ll need it as an environment variable (ELEVENLABS_API_KEY) or directly in the script (not recommended for production).

Setting up the project

Create a fresh directory and initialise a virtual environment:

mkdir audiobook-generator && cd audiobook-generator
python -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install requests tqdm
Enter fullscreen mode Exit fullscreen mode

requests handles HTTP calls, while tqdm gives us a nice progress bar when processing many chapters.

Create a simple folder layout:

audiobook-generator/
│
├─ chapters/        # plain‑text files, one per chapter (chapter1.txt, chapter2.txt, …)
├─ output/          # generated MP3s will land here
└─ generate_audiobook.py
Enter fullscreen mode Exit fullscreen mode

Place your source text files in chapters/. For the demo we’ll assume each file contains a single chapter of a public‑domain novel.


Converting text to speech with ElevenLabs

ElevenLabs’ TTS endpoint expects a JSON payload with the text, voice ID, and optional settings (stability, similarity boost, etc.). Below is a minimal Python wrapper that:

  • Reads a chapter file
  • Calls the API with streaming enabled (so we can write the MP3 directly to disk)
  • Saves the result as output/chapterX.mp3
# generate_audiobook.py
import os
import json
import requests
from tqdm import tqdm

API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
VOICE_ID = "EXAVITQu4vr4xnSDxMaL"   # default “Rachel” voice – replace with your own

def tts(text: str, outfile: str):
    url = f"{BASE_URL}/text-to-speech/{VOICE_ID}"
    headers = {
        "xi-api-key": API_KEY,
        "Content-Type": "application/json",
        "Accept": "audio/mpeg"
    }
    payload = {
        "text": text,
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {
            "stability": 0.75,
            "similarity_boost": 0.75
        }
    }

    # Stream the binary audio back to disk
    with requests.post(url, headers=headers, json=payload, stream=True) as r:
        r.raise_for_status()
        with open(outfile, "wb") as f:
            for chunk in r.iter_content(chunk_size=8192):
                f.write(chunk)

def main():
    os.makedirs("output", exist_ok=True)
    chapter_files = sorted([f for f in os.listdir("chapters") if f.endswith(".txt")])

    for idx, fname in enumerate(tqdm(chapter_files, desc="Generating audio")):
        with open(os.path.join("chapters", fname), "r", encoding="utf-8") as f:
            text = f.read()
        out_path = os.path.join("output", f"chapter{idx+1}.mp3")
        tts(text, out_path)

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Quick curl alternative

If you prefer a one‑liner for a single chapter, here’s how you’d do it with curl:

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \
  -H "xi-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Accept: audio/mpeg" \
  -d '{
        "text": "Once upon a time, in a far‑away kingdom...",
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {"stability":0.75,"similarity_boost":0.75}
      }' \
  --output chapter1.mp3
Enter fullscreen mode Exit fullscreen mode

Replace YOUR_API_KEY with the key you grabbed earlier. This is handy for quick tests or for integrating into shell scripts.


Stitching audio into an audiobook

Now that each chapter lives as an individual MP3, we’ll concatenate them into a single file. ffmpeg can do this without re‑encoding, preserving the original quality.

# Inside the project root
cd output

# Create a temporary file list for ffmpeg
for f in chapter*.mp3; do echo "file '$PWD/$f'" >> files.txt; done

# Concatenate
ffmpeg -f concat -safe 0 -i files.txt -c copy ../audiobook.mp3

# Clean up
rm files.txt
Enter fullscreen mode Exit fullscreen mode

The resulting audiobook.mp3 sits in the project root and is ready for distribution.


Bonus: Adding chapter metadata

Many audiobook players respect ID3 tags. You can inject chapter titles, authors, and even cover art using mutagen (a pure‑Python ID3 library).

pip install mutagen
Enter fullscreen mode Exit fullscreen mode
# add_metadata.py
from mutagen.id3 import ID3, TIT2, TALB, TPE1, APIC
from mutagen.mp3 import MP3

def tag_audiobook(file_path, title, author, album, cover_path=None):
    audio = MP3(file_path, ID3=ID3)

    # Add basic tags
    audio["TIT2"] = TIT2(encoding=3, text=title)   # Title
    audio["TPE1"] = TPE1(encoding=3, text=author)  # Artist/Author
    audio["TALB"] = TALB(encoding=3, text=album)   # Album (e.g., book title)

    # Optional cover image
    if cover_path and os.path.isfile(cover_path):
        with open(cover_path, "rb") as img:
            audio["APIC"] = APIC(
                encoding=3,
                mime="image/jpeg",
                type=3,
                desc="Cover",
                data=img.read()
            )
    audio.save()

# Example usage
tag_audiobook(
    "audiobook.mp3",
    title="The Adventures of Sherlock Holmes",
    author="Arthur Conan Doyle",
    album="Sherlock Holmes Collection",
    cover_path="cover.jpg"
)
Enter fullscreen mode Exit fullscreen mode

Run this after you generate audiobook.mp3 to embed proper metadata, making the file look professional in any player.


Wrap‑up

You now have a fully functional pipeline:

  1. Fetch or write plain‑text chapters → chapters/
  2. Call ElevenLabs to synthesize each chunk → output/chapterX.mp3
  3. Combine the MP3s with ffmpeg → audiobook.mp3
  4. (Optional) Tag the final file for a polished listening experience

Because the ElevenLabs API supports voice cloning, you can even train a custom voice from your own recordings and use it to narrate the entire book in a truly unique timbre. The same code works for podcasts, e‑learning modules, or any scenario where you need high‑quality TTS on the fly.


Next steps

  • Experiment with different voice IDs (ElevenLabs offers several pre‑built voices).
  • Play with stability and similarity_boost to fine‑tune prosody for longer passages.
  • Integrate the script into a CI/CD job so every new chapter you push to a repo automatically becomes part of the audiobook.

If you found this guide helpful, give it a spin and see how quickly you can turn text into an immersive listening experience.

Ready to generate your first audiobook? Grab your ElevenLabs API key and start building at https://try.elevenlabs.io/kr07zfuqn1bp. Happy coding!

Top comments (0)