DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Create an Interactive Audio Story Generator

Introduction

Hey devs!

Have you ever wanted to turn a simple script into a living, breathing audio experience? Imagine a choose‑your‑own adventure where the narrator’s voice changes based on the player’s choices, or a podcast that adapts on the fly to the listener’s mood. That’s the power of an Interactive Audio Story Generator (IASG). In this post we’ll walk through building one from scratch, leveraging modern text‑to‑speech (TTS) and voice‑cloning tech. And if you’re curious about the best tool for the job, keep an eye on the section about ElevenLabs—trust me, you’ll want to try it out.

Why Interactive Audio Stories?

  1. Accessibility: Audio narratives reach users who can’t read easily or prefer listening.
  2. Immersion: Dynamic narration feels like a conversation with the story itself.
  3. Scalability: Once the TTS pipeline is set up, you can generate thousands of unique episodes without a human narrator.

The core of any IASG is a voice engine that can produce natural, expressive speech from text. That’s where TTS and voice cloning come in.

Setting Up the Project

Let’s start with a lightweight stack:

  • Python for the backend TTS orchestration.
  • FastAPI to expose a simple HTTP endpoint.
  • WebSocket (via fastapi-websocket) to stream audio chunks to the front end.
  • React on the client for UI and audio playback.

You can bootstrap this with:

pip install fastapi uvicorn python-multipart websockets
npm create vite@latest ia-story -- --template react
Enter fullscreen mode Exit fullscreen mode

Text‑to‑Speech with ElevenLabs

ElevenLabs offers a developer‑friendly TTS API with high‑fidelity voice models and voice cloning capabilities. Their pricing is straightforward and the SDKs are well documented. Here’s how to get started in Python:

import requests

ELEVENLABS_API_KEY = "YOUR_API_KEY"
BASE_URL = "https://api.elevenlabs.io/v1"

def synthesize(text, voice_id="en-US-Standard-A"):
    headers = {
        "xi-api-key": ELEVENLABS_API_KEY,
        "Content-Type": "application/json",
    }
    payload = {
        "text": text,
        "voice_id": voice_id,
        "model_id": "eleven_monolingual_v1",
    }
    resp = requests.post(f"{BASE_URL}/text-to-speech/{voice_id}", json=payload, headers=headers)
    resp.raise_for_status()
    return resp.content  # raw audio bytes
Enter fullscreen mode Exit fullscreen mode

Pro tip: Use the eleven_monolingual_v1 model for the best balance of speed and quality.

You can find the official SDK at the same link: https://try.elevenlabs.io/kr07zfuqn1bp

Voice Cloning 101

If you want the narrator to sound like a particular character or a real person, ElevenLabs supports voice cloning. The process is:

  1. Upload a short audio sample (30‑60 seconds).
  2. Let the model generate a voice ID.
  3. Use that voice ID in subsequent TTS requests.
curl -X POST "https://api.elevenlabs.io/v1/voices" \
     -H "xi-api-key: $ELEVENLABS_API_KEY" \
     -F "audio=@/path/to/sample.wav"
Enter fullscreen mode Exit fullscreen mode

The response will include a voice_id you can reuse.

Heads‑up: Always respect privacy and copyright when cloning voices.

Building the Interactive Flow

An IASG needs to decide what to say next based on user input. Here’s a minimal state machine:

from fastapi import FastAPI, WebSocket
app = FastAPI()

story_nodes = {
    "start": {
        "text": "You find yourself in a dim forest. Do you go left or right?",
        "choices": {"left": "left_node", "right": "right_node"}
    },
    "left_node": {"text": "A river blocks your path.", "choices": {}},
    "right_node": {"text": "A clearing appears with a mysterious hut.", "choices": {}},
}

@app.websocket("/ws/story")
async def story_ws(ws: WebSocket):
    await ws.accept()
    current = "start"
    while True:
        node = story_nodes[current]
        audio = synthesize(node["text"])
        await ws.send_bytes(audio)  # stream audio to client
        if not node["choices"]:
            break
        # wait for client choice
        data = await ws.receive_text()
        current = node["choices"].get(data, "start")
Enter fullscreen mode Exit fullscreen mode

The client would play the received audio bytes and provide a simple UI (buttons) to send the next choice back through the WebSocket.

Sample Front‑End (React)

import { useEffect, useRef, useState } from "react";

export default function StoryPlayer() {
  const [audioURL, setAudioURL] = useState<string>("");
  const ws = useRef<WebSocket | null>(null);

  useEffect(() => {
    ws.current = new WebSocket("ws://localhost:8000/ws/story");
    ws.current.binaryType = "arraybuffer";

    ws.current.onmessage = (e) => {
      const blob = new Blob([e.data], { type: "audio/mpeg" });
      setAudioURL(URL.createObjectURL(blob));
    };

    return () => ws.current?.close();
  }, []);

  const sendChoice = (choice: string) => {
    ws.current?.send(choice);
  };

  return (
    <div>
      {audioURL && <audio src={audioURL} controls autoPlay />}
      <button onClick={() => sendChoice("left")}>Left</button>
      <button onClick={() => sendChoice("right")}>Right</button>
    </div>
  );
}
Enter fullscreen mode Exit fullscreen mode

Tips & Best Practices

Tip Why it matters
Chunk your audio Streaming smaller chunks reduces latency.
Cache common nodes Re‑synthesize rarely‑changed text only once.
Use SSML Add pauses, emphasis, and pitch changes for naturalness.
Rate‑limit ElevenLabs enforces limits; back‑off on 429 responses.

Deploying

For production, consider:

  • Docker to bundle FastAPI and dependencies.
  • NGINX or Caddy for TLS termination.
  • Cloud Run / ECS for serverless scaling.
  • WebSocket support: ensure the host forwards Upgrade headers.
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
Enter fullscreen mode Exit fullscreen mode

Conclusion

Building an Interactive Audio Story Generator is a fun way to blend storytelling with cutting‑edge voice tech. By combining FastAPI, WebSockets, and a powerful TTS engine like ElevenLabs, you can deliver a truly immersive, dynamic audio experience. And if you haven’t already, give ElevenLabs a spin—they make voice synthesis feel like magic.

Ready to bring your stories to life?

Try ElevenLabs today and start generating stunning audio narratives: https://try.elevenlabs.io/kr07zfuqn1bp

Top comments (0)