Introduction
Ever wanted to turn a novel, a blog series, or a collection of technical docs into a polished audiobook—without spending hours in a recording studio? With modern voice‑AI services you can generate natural‑sounding narration programmatically, and ElevenLabs makes it surprisingly easy. In this walkthrough we’ll build a simple “Audiobook Generator” that pulls plain‑text chapters, sends them to ElevenLabs’ Text‑to‑Speech (TTS) API, and stitches the resulting MP3 files together into a single audiobook file.
You’ll walk away with a reusable Python script (and a quick curl example) that you can plug into any content pipeline, whether you’re a solo indie author, a SaaS that serves audio versions of articles, or just a hobbyist who loves to experiment with voice cloning.
Prerequisites
| What you need | Why |
|---|---|
| Python 3.8+ (or Node if you prefer) | To drive the API calls and file handling |
| An ElevenLabs account (free tier works for short clips) | Provides the TTS endpoint and voice models |
ffmpeg installed and on your PATH
|
Concatenates MP3 segments into a single file |
| Basic familiarity with REST APIs | Helps you tweak request parameters |
If you don’t have ffmpeg yet, install it via your package manager:
# macOS
brew install ffmpeg
# Ubuntu/Debian
sudo apt-get install ffmpeg
Getting an API key
- Sign up at the ElevenLabs platform using this affiliate link: https://try.elevenlabs.io/kr07zfuqn1bp.
- Once logged in, navigate to API → Keys and generate a new key.
- Copy the key; you’ll need it as an environment variable (
ELEVENLABS_API_KEY) or directly in the script (not recommended for production).
Setting up the project
Create a fresh directory and initialise a virtual environment:
mkdir audiobook-generator && cd audiobook-generator
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install requests tqdm
requests handles HTTP calls, while tqdm gives us a nice progress bar when processing many chapters.
Create a simple folder layout:
audiobook-generator/
│
├─ chapters/ # plain‑text files, one per chapter (chapter1.txt, chapter2.txt, …)
├─ output/ # generated MP3s will land here
└─ generate_audiobook.py
Place your source text files in chapters/. For the demo we’ll assume each file contains a single chapter of a public‑domain novel.
Converting text to speech with ElevenLabs
ElevenLabs’ TTS endpoint expects a JSON payload with the text, voice ID, and optional settings (stability, similarity boost, etc.). Below is a minimal Python wrapper that:
- Reads a chapter file
- Calls the API with streaming enabled (so we can write the MP3 directly to disk)
- Saves the result as
output/chapterX.mp3
# generate_audiobook.py
import os
import json
import requests
from tqdm import tqdm
API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
VOICE_ID = "EXAVITQu4vr4xnSDxMaL" # default “Rachel” voice – replace with your own
def tts(text: str, outfile: str):
url = f"{BASE_URL}/text-to-speech/{VOICE_ID}"
headers = {
"xi-api-key": API_KEY,
"Content-Type": "application/json",
"Accept": "audio/mpeg"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.75
}
}
# Stream the binary audio back to disk
with requests.post(url, headers=headers, json=payload, stream=True) as r:
r.raise_for_status()
with open(outfile, "wb") as f:
for chunk in r.iter_content(chunk_size=8192):
f.write(chunk)
def main():
os.makedirs("output", exist_ok=True)
chapter_files = sorted([f for f in os.listdir("chapters") if f.endswith(".txt")])
for idx, fname in enumerate(tqdm(chapter_files, desc="Generating audio")):
with open(os.path.join("chapters", fname), "r", encoding="utf-8") as f:
text = f.read()
out_path = os.path.join("output", f"chapter{idx+1}.mp3")
tts(text, out_path)
if __name__ == "__main__":
main()
Quick curl alternative
If you prefer a one‑liner for a single chapter, here’s how you’d do it with curl:
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \
-H "xi-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: audio/mpeg" \
-d '{
"text": "Once upon a time, in a far‑away kingdom...",
"model_id": "eleven_monolingual_v1",
"voice_settings": {"stability":0.75,"similarity_boost":0.75}
}' \
--output chapter1.mp3
Replace YOUR_API_KEY with the key you grabbed earlier. This is handy for quick tests or for integrating into shell scripts.
Stitching audio into an audiobook
Now that each chapter lives as an individual MP3, we’ll concatenate them into a single file. ffmpeg can do this without re‑encoding, preserving the original quality.
# Inside the project root
cd output
# Create a temporary file list for ffmpeg
for f in chapter*.mp3; do echo "file '$PWD/$f'" >> files.txt; done
# Concatenate
ffmpeg -f concat -safe 0 -i files.txt -c copy ../audiobook.mp3
# Clean up
rm files.txt
The resulting audiobook.mp3 sits in the project root and is ready for distribution.
Bonus: Adding chapter metadata
Many audiobook players respect ID3 tags. You can inject chapter titles, authors, and even cover art using mutagen (a pure‑Python ID3 library).
pip install mutagen
# add_metadata.py
from mutagen.id3 import ID3, TIT2, TALB, TPE1, APIC
from mutagen.mp3 import MP3
def tag_audiobook(file_path, title, author, album, cover_path=None):
audio = MP3(file_path, ID3=ID3)
# Add basic tags
audio["TIT2"] = TIT2(encoding=3, text=title) # Title
audio["TPE1"] = TPE1(encoding=3, text=author) # Artist/Author
audio["TALB"] = TALB(encoding=3, text=album) # Album (e.g., book title)
# Optional cover image
if cover_path and os.path.isfile(cover_path):
with open(cover_path, "rb") as img:
audio["APIC"] = APIC(
encoding=3,
mime="image/jpeg",
type=3,
desc="Cover",
data=img.read()
)
audio.save()
# Example usage
tag_audiobook(
"audiobook.mp3",
title="The Adventures of Sherlock Holmes",
author="Arthur Conan Doyle",
album="Sherlock Holmes Collection",
cover_path="cover.jpg"
)
Run this after you generate audiobook.mp3 to embed proper metadata, making the file look professional in any player.
Wrap‑up
You now have a fully functional pipeline:
-
Fetch or write plain‑text chapters →
chapters/ -
Call ElevenLabs to synthesize each chunk →
output/chapterX.mp3 -
Combine the MP3s with
ffmpeg→audiobook.mp3 - (Optional) Tag the final file for a polished listening experience
Because the ElevenLabs API supports voice cloning, you can even train a custom voice from your own recordings and use it to narrate the entire book in a truly unique timbre. The same code works for podcasts, e‑learning modules, or any scenario where you need high‑quality TTS on the fly.
Next steps
- Experiment with different voice IDs (ElevenLabs offers several pre‑built voices).
- Play with
stabilityandsimilarity_boostto fine‑tune prosody for longer passages. - Integrate the script into a CI/CD job so every new chapter you push to a repo automatically becomes part of the audiobook.
If you found this guide helpful, give it a spin and see how quickly you can turn text into an immersive listening experience.
Ready to generate your first audiobook? Grab your ElevenLabs API key and start building at https://try.elevenlabs.io/kr07zfuqn1bp. Happy coding!
Top comments (0)