DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Build a Voice-Powered Accessibility Tool

Why Voice‑Powered Accessibility Matters

If you’ve ever struggled to read a long article on a tiny screen, you know how powerful a good voice interface can be. For people with visual impairments, dyslexia, or motor challenges, turning text into natural‑sounding speech isn’t just a convenience—it’s a necessity. As developers, we have the tools to build solutions that make the web and apps more inclusive, and the best part is that modern APIs let us do it with just a few lines of code.

In this article we’ll walk through a practical, end‑to‑end example of a voice‑powered accessibility tool:

  1. Convert arbitrary text to speech (TTS) using a high‑quality service.
  2. Clone a custom voice so the output feels personal and consistent.
  3. Hook everything up to a tiny web UI that anyone can use.

All of this is powered by ElevenLabs, a TTS platform that delivers ultra‑realistic voices and a straightforward REST API. You can sign up for free and get an API key right away: https://try.elevenlabs.io/kr07zfuqn1bp


Getting Your ElevenLabs API Key

  1. Visit the signup page (the link above) and create an account.
  2. Once logged in, navigate to Dashboard → API Keys and generate a new key.
  3. Keep this key safe; you’ll need it for every request you make.

Tip: Store the key in an environment variable (ELEVENLABS_API_KEY) instead of hard‑coding it. This keeps your credentials out of source control.


Quick TTS Test with curl

Before we dive into code, let’s verify that the API works. Open a terminal and run:

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/voice-id" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "text": "Hello, world! This is a quick test of ElevenLabs TTS.",
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {"stability": 0.75, "similarity_boost": 0.85}
      }' \
  --output hello.wav
Enter fullscreen mode Exit fullscreen mode

If everything is set up correctly, you’ll get a hello.wav file that sounds surprisingly human. Replace "voice-id" with the ID of the default voice you want to use (you can list available voices via the API or the dashboard).


Building a Python Wrapper

Most developers prefer to work in Python for quick prototyping. Below is a tiny wrapper that abstracts the HTTP call and returns an in‑memory audio buffer.

import os
import requests
from io import BytesIO

ELEVENLABS_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"

def synthesize(text: str, voice_id: str, stability=0.75, similarity=0.85):
    url = f"{BASE_URL}/text-to-speech/{voice_id}"
    headers = {
        "xi-api-key": ELEVENLABS_API_KEY,
        "Content-Type": "application/json"
    }
    payload = {
        "text": text,
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {
            "stability": stability,
            "similarity_boost": similarity
        }
    }
    resp = requests.post(url, json=payload, headers=headers)
    resp.raise_for_status()
    return BytesIO(resp.content)   # returns a file‑like object

# Example usage
if __name__ == "__main__":
    voice_id = "EXAVITQu4vr4xnSDxMaL"  # replace with your chosen voice
    audio = synthesize("Accessibility matters for everyone.", voice_id)
    with open("output.wav", "wb") as f:
        f.write(audio.read())
Enter fullscreen mode Exit fullscreen mode

Save this as tts.py and run python tts.py. You should see an output.wav file ready for playback.


Adding Voice Cloning

ElevenLabs shines when you need a custom voice—for example, a brand mascot or a user‑specific voice that makes the experience feel personal. The cloning workflow is:

  1. Upload a short (≈30‑second) voice sample.
  2. Wait for the service to process it (usually < 2 minutes).
  3. Use the returned voice_id in subsequent TTS calls.

Here’s a Python snippet that uploads a sample and fetches the new voice ID:

def clone_voice(name: str, audio_path: str):
    url = f"{BASE_URL}/voices/add"
    headers = {"xi-api-key": ELEVENLABS_API_KEY}
    files = {
        "name": (None, name),
        "files": (os.path.basename(audio_path), open(audio_path, "rb"), "audio/wav")
    }
    resp = requests.post(url, files=files, headers=headers)
    resp.raise_for_status()
    voice_data = resp.json()
    return voice_data["voice_id"]

# Usage
my_voice_id = clone_voice("MyAssistant", "my_sample.wav")
print(f"Cloned voice ID: {my_voice_id}")
Enter fullscreen mode Exit fullscreen mode

Once you have my_voice_id, simply pass it to the synthesize function from the previous section. The result will be spoken in the cloned voice, preserving the speaker’s unique timbre.

Pro tip: Keep the sample clean (no background noise) and use a consistent speaking rate. The better the source, the more natural the cloned voice will sound.


Front‑End Integration with JavaScript

A truly accessible tool needs a UI that anyone can reach. Let’s create a minimal HTML page that lets users paste text, hit “Speak”, and hear the result instantly. We’ll use the browser’s fetch API to call a small backend endpoint (the Python server we just wrote) and then play the returned audio with the Web Audio API.

index.html

<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <title>Voice Accessibility Demo</title>
  <style>
    body { font-family: sans-serif; margin: 2rem; }
    textarea { width: 100%; height: 120px; }
    button { padding: .5rem 1rem; margin-top: .5rem; }
  </style>
</head>
<body>
  <h1>Speak Your Text</h1>
  <textarea id="txt" placeholder="Enter something to read..."></textarea><br/>
  <button id="speakBtn">Speak</button>

  <script>
    const btn = document.getElementById('speakBtn');
    const txt = document.getElementById('txt');

    btn.onclick = async () => {
      const response = await fetch('/speak', {
        method: 'POST',
        headers: { 'Content-Type': 'application/json' },
        body: JSON.stringify({ text: txt.value })
      });

      if (!response.ok) {
        alert('Error: ' + response.statusText);
        return;
      }

      const arrayBuffer = await response.arrayBuffer();
      const audioContext = new (window.AudioContext || window.webkitAudioContext)();
      const audioBuffer = await audioContext.decodeAudioData(arrayBuffer);
      const source = audioContext.createBufferSource();
      source.buffer = audioBuffer;
      source.connect(audioContext.destination);
      source.start(0);
    };
  </script>
</body>
</html>
Enter fullscreen mode Exit fullscreen mode

Backend (app.py) – Flask example

from flask import Flask, request, send_file
from tts import synthesize   # the function from earlier
import os
import io

app = Flask(__name__)

DEFAULT_VOICE = os.getenv("DEFAULT_VOICE_ID", "EXAVITQu4vr4xnSDxMaL")

@app.route("/speak", methods=["POST"])
def speak():
    data = request.get_json()
    text = data.get("text", "")
    audio_io = synthesize(text, DEFAULT_VOICE)
    audio_io.seek(0)
    return send_file(
        audio_io,
        mimetype="audio/wav",
        as_attachment=False,
        download_name="speech.wav"
    )

if __name__ == "__main__":
    app.run(debug=True)
Enter fullscreen mode Exit fullscreen mode

Run python app.py, open http://localhost:5000 in a browser, type something, and press Speak. The text is sent to the Flask route, which calls ElevenLabs behind the scenes, and the resulting audio streams back to the client for immediate playback.


Making It Truly Accessible

Now that the core pipeline works, consider these enhancements to meet accessibility guidelines:

Feature Why It Helps Quick Implementation
Adjustable speaking rate Users with cognitive impairments may need slower speech. Pass speed (or stability) values to the API payload.
Language selection Non‑English speakers need native‑language voices. ElevenLabs supports multiple languages; swap model_id accordingly.
Keyboard shortcuts Reduces reliance on mouse navigation. Add keydown listeners for Enter or Space to trigger the button.
Downloadable audio Some users prefer an offline copy. Add a “Download” link that points to the same /speak endpoint with Content-Disposition: attachment.

Each of these tweaks can be added in a handful of lines, but they dramatically increase the tool’s reach.


Scaling Up

If you plan to serve many users, keep these production concerns in mind:

  • Rate limits – ElevenLabs enforces per‑minute limits based on your plan. Cache repeated requests or batch them when possible.
  • Secure API keys – Never expose the key to the browser. All calls to ElevenLabs must go through a server you control.
  • Cost monitoring – TTS is billed per generated character. Log usage and set alerts to avoid surprise bills.

ElevenLabs offers generous free tiers for developers, but you’ll want to upgrade if you cross the threshold of a few hundred thousand characters per month. More details are in the dashboard.


Next Steps

  1. Experiment with voice cloning – Record a short phrase in your own voice and see how the system mimics you.
  2. Add multi‑modal input – Combine speech‑to‑text (STT) with the TTS pipeline for a full‑duplex voice assistant.
  3. Publish as an npm package – Wrap the HTTP calls into a reusable library for the broader community.

Ready to Give It a Spin?

If you’re excited to bring natural, human‑like speech to your apps, the fastest way to start is by signing up for ElevenLabs. The platform’s API is clean, the voices are top‑tier, and you’ll get a generous free quota to prototype without friction. Grab your API key now and build something that truly makes the web accessible for everyone: https://try.elevenlabs.io/kr07zfuqn1bp

Happy coding, and may your applications speak as clearly as you think!

Top comments (0)