DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

How to Add Voice to Your React App with ElevenLabs

Why Voice Matters in Modern Web Apps

If you’re building a chatbot, a language‑learning tool, or a virtual tour guide, you’ll quickly discover that text‑only interactions feel flat. Adding spoken output turns a static UI into an engaging, accessible experience that can reach users who prefer listening over reading. The technology behind this—text‑to‑speech (TTS)—has come a long way, and today it’s surprisingly straightforward to hook a high‑quality voice engine into a React project.

In this post we’ll walk through:

  • Why TTS is a game‑changer for web developers
  • The key concepts you need to understand (voices, prompts, SSML)
  • How to integrate ElevenLabs, a leading TTS provider, into a React app
  • A practical example that includes voice cloning
  • Tips for optimizing performance and user experience

Let’s dive in!


The Voice Advantage

  1. Accessibility

    Screen‑reader users benefit from native TTS, but a custom voice lets you tailor the tone, pace, and personality of the narration.

  2. Engagement

    Auditory feedback keeps users on task—think of product walkthroughs, e‑learning modules, or interactive storytelling.

  3. Multilingual Reach

    With AI voices that support dozens of languages, you can serve a global audience without a multilingual team.

  4. Brand Personality

    A consistent voice across your app reinforces brand identity, just like a logo or color palette.


Core TTS Concepts

Concept What it is Why it matters
Voice A pre‑trained neural model that defines gender, accent, age, and style. Determines how the spoken output feels.
Prompt The raw text you want to convert to speech. The content that the model reads.
SSML Speech Synthesis Markup Language. Gives you control over pauses, emphasis, pitch, and more. Lets you fine‑tune the narration for naturalness.
Voice Cloning Creating a custom voice model from a few minutes of audio. Enables a brand‑specific voice or a personal assistant.

Choosing a Provider: ElevenLabs

While there are many TTS APIs out there, ElevenLabs consistently delivers the most natural‑sounding voices and a developer‑friendly SDK. They support voice cloning, advanced SSML, and low‑latency streaming—all essential for a smooth React experience.

Try ElevenLabs today: https://try.elevenlabs.io/kr07zfuqn1bp


Setting Up the Project

First, create a new React app (or use an existing one). We’ll use Vite for its lightning speed, but the same approach works with Create‑React‑App or Next.js.

npm create vite@latest tts-demo -- --template react
cd tts-demo
npm install
Enter fullscreen mode Exit fullscreen mode

Install the ElevenLabs SDK

ElevenLabs ships a lightweight JavaScript client that handles authentication and request formatting.

npm install @elevenlabs/client
Enter fullscreen mode Exit fullscreen mode

Tip: Store your API key in an .env file and load it with Vite’s import.meta.env.

# .env
VITE_ELEVENLABS_API_KEY=sk_your_key_here
Enter fullscreen mode Exit fullscreen mode
// src/api/elevenLabs.js
import ElevenLabs from '@elevenlabs/client';

export const elevenClient = new ElevenLabs({
  apiKey: import.meta.env.VITE_ELEVENLABS_API_KEY,
});
Enter fullscreen mode Exit fullscreen mode

Basic Text‑to‑Speech

Let’s start with a simple component that takes user input and plays the spoken version.

// src/components/TextToSpeech.jsx
import { useState } from 'react';
import { elevenClient } from '../api/elevenLabs';

export default function TextToSpeech() {
  const [text, setText] = useState('');
  const [isSpeaking, setIsSpeaking] = useState(false);

  const handleSpeak = async () => {
    setIsSpeaking(true);
    try {
      const voiceId = '21m00Tcm4TlvDq8ikWAM'; // Default ElevenLabs voice
      const { audioUrl } = await elevenClient.textToSpeech({
        voice: voiceId,
        text,
        outputFormat: 'mp3',
      });

      const audio = new Audio(audioUrl);
      audio.onended = () => setIsSpeaking(false);
      audio.play();
    } catch (error) {
      console.error('TTS error:', error);
      setIsSpeaking(false);
    }
  };

  return (
    <div className="p-4 max-w-md mx-auto">
      <textarea
        className="w-full h-32 p-2 border rounded"
        placeholder="Enter text to speak…"
        value={text}
        onChange={(e) => setText(e.target.value)}
      />
      <button
        className="mt-2 px-4 py-2 bg-blue-600 text-white rounded disabled:opacity-50"
        onClick={handleSpeak}
        disabled={isSpeaking || !text}
      >
        {isSpeaking ? 'Speaking…' : 'Speak'}
      </button>
    </div>
  );
}
Enter fullscreen mode Exit fullscreen mode

What’s Happening Here?

  1. Voice ID – We’re using ElevenLabs’ default voice; you can pick any from their catalog.
  2. Output Format – mp3 is widely supported; you can also request wav or ogg.
  3. Streaming – The SDK returns a direct URL to the synthesized audio, which we play with the browser’s native <audio> element.

Adding SSML for Naturalness

SSML lets you control pacing, emphasis, and pauses. ElevenLabs accepts SSML strings directly.

const ssml = `
  <speak>
    <voice name="21m00Tcm4TlvDq8ikWAM">
      Hello, <break time="500ms"/> welcome to our <emphasis level="strong">voice demo</emphasis>.
    </voice>
  </speak>
`;

const { audioUrl } = await elevenClient.textToSpeech({
  voice: '21m00Tcm4TlvDq8ikWAM',
  text: ssml,
  outputFormat: 'mp3',
  voiceSettings: { stability: 0.5, similarityBoost: 0.8 },
  useSSML: true,
});
Enter fullscreen mode Exit fullscreen mode

Pro tip: Keep SSML concise; the API can handle it in under a second for short prompts.


Voice Cloning: Create Your Own Voice

ElevenLabs’ voice cloning lets you generate a custom voice model from a few minutes of recorded speech. This is perfect for brand consistency or personal assistants.

Step 1 – Upload Audio

curl -X POST "https://api.elevenlabs.io/v1/voices" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" \
  -H "Content-Type: multipart/form-data" \
  -F "file=@/path/to/your/audio.wav" \
  -F "name=MyCustomVoice"
Enter fullscreen mode Exit fullscreen mode

The response will include a voice_id you’ll use for synthesis.

Step 2 – Synthesize with Your Custom Voice

const { audioUrl } = await elevenClient.textToSpeech({
  voice: 'MyCustomVoice', // ID returned from the upload
  text: 'Hello from my brand‑specific voice!',
  outputFormat: 'mp3',
});
Enter fullscreen mode Exit fullscreen mode

Heads‑up: Cloning requires a high‑quality audio file (at least 5 minutes) and may take a few minutes to process.


Optimizing Performance

  1. Lazy Load Audio

    Load the ElevenLabs SDK only when the user clicks “Speak” to reduce initial bundle size.

  2. Cache Audio

    Store previously synthesized audio URLs in localStorage or a service worker cache to avoid re‑fetching.

  3. Use Web Workers

    For heavy SSML processing or large prompts, offload the API call to a worker to keep the UI responsive.

  4. Graceful Fallback

    If the network fails, display a friendly error message and suggest trying again later.


Accessibility Checklist

  • Keyboard navigation – Ensure buttons are reachable via Tab and can be activated with Enter/Space.
  • ARIA labels – Add aria-live="polite" to the status area so screen readers announce “Speaking…”.
  • Volume controls – Provide a volume slider or mute button for users who need to adjust audio levels.

Next Steps

  1. Add Voice to Your Forms – Announce form validation errors or success messages aloud.
  2. Create a Conversational UI – Combine TTS with speech‑to‑text for a hands‑free chatbot.
  3. Personalize with SSML – Use user‑specific data (e.g., name, preferences) to tailor the narration.

Call to Action

Ready to bring your React app to life with natural, expressive voices? ElevenLabs offers a free tier and seamless integration for developers. Sign up today and start building engaging, voice‑rich experiences with minimal effort.

Try ElevenLabs now: https://try.elevenlabs.io/kr07zfuqn1bp

Happy coding—and happy listening!

Top comments (0)