DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Create an AI Voice Assistant with ElevenLabs and Node.js

Why ElevenLabs is a Game‑Changer for Voice AI

If you’ve ever tried to build a voice assistant, you know the biggest headache is getting natural‑sounding speech. Traditional TTS engines either sound robotic or require massive amounts of data and compute. ElevenLabs solves both problems with a cloud API that delivers studio‑grade speech and even lets you clone a voice in minutes. The best part? You can start using it from a simple Node.js script and scale up to full‑blown conversational agents.

Quick tip: Sign up through this affiliate link – it gives you a free credit to experiment: https://try.elevenlabs.io/kr07zfuqn1bp

Overview of the Architecture

At a high level, our AI voice assistant will consist of three parts:

  1. Speech‑to‑Text (STT) – Convert the user’s microphone input to text. For this demo we’ll use the free Web Speech API in the browser.
  2. Logic Layer – A small Node.js server that receives the transcribed text, decides on a response (you can plug in OpenAI, Cohere, or any LLM), and asks ElevenLabs to synthesize the reply.
  3. Text‑to‑Speech (TTS) – The ElevenLabs API turns the response text into an audio buffer that we stream back to the client.

All the heavy lifting is done by ElevenLabs, so you can focus on the conversational logic.

Setting Up the Project

# Create a fresh folder
mkdir ai-voice-assistant && cd ai-voice-assistant

# Initialize a Node project
npm init -y

# Install dependencies
npm install express axios cors dotenv
Enter fullscreen mode Exit fullscreen mode

Create a .env file to keep your API key safe:

ELEVENLABS_API_KEY=your_elevenlabs_api_key_here
PORT=3000
Enter fullscreen mode Exit fullscreen mode

Note: Grab your API key from the ElevenLabs dashboard after signing up via https://try.elevenlabs.io/kr07zfuqn1bp.

The Express Server

Below is a minimal server that receives a JSON payload like { "text": "Hello, world!" } and returns a base64‑encoded audio string.

// server.js
require('dotenv').config();
const express = require('express');
const axios = require('axios');
const cors = require('cors');

const app = express();
app.use(cors());
app.use(express.json());

const ELEVENLABS_API_KEY = process.env.ELEVENLABS_API_KEY;
const VOICE_ID = 'EXAVITQu4vr4xnSDxMaL'; // default voice; replace with your cloned voice ID

app.post('/synthesize', async (req, res) => {
  const { text } = req.body;
  if (!text) return res.status(400).json({ error: 'Missing text' });

  try {
    const response = await axios({
      method: 'post',
      url: `https://api.elevenlabs.io/v1/text-to-speech/${VOICE_ID}`,
      headers: {
        'xi-api-key': ELEVENLABS_API_KEY,
        'Content-Type': 'application/json',
        'Accept': 'audio/mpeg',
      },
      data: {
        text,
        voice_settings: {
          stability: 0.75,
          similarity_boost: 0.85,
        },
      },
      responseType: 'arraybuffer',
    });

    const base64Audio = Buffer.from(response.data, 'binary').toString('base64');
    res.json({ audio: base64Audio });
  } catch (err) {
    console.error('ElevenLabs error:', err.response?.data || err.message);
    res.status(500).json({ error: 'TTS failed' });
  }
});

app.listen(process.env.PORT, () => {
  console.log(`🚀 Server listening on http://localhost:${process.env.PORT}`);
});
Enter fullscreen mode Exit fullscreen mode

What’s happening here?

  • Voice ID – You can use any of ElevenLabs’ premade voices, or the ID of a voice you’ve cloned (more on that later).
  • Stability & Similarity Boost – Tweaking these numbers changes how “steady” the voice sounds and how closely it matches the reference voice.
  • Response Type – We ask for audio/mpeg (MP3) because it streams easily in browsers.

Run the server:

node server.js
Enter fullscreen mode Exit fullscreen mode

Front‑End: Capturing Speech and Playing Back Audio

Create a simple HTML page that uses the Web Speech API for STT and fetches the synthesized audio from our server.

<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <title>AI Voice Assistant Demo</title>
  <style>
    body { font-family: sans-serif; margin: 2rem; }
    button { padding: .5rem 1rem; margin-top: 1rem; }
  </style>
</head>
<body>
  <h1>Talk to the AI</h1>
  <button id="talkBtn">Start Listening</button>
  <p id="transcript"></p>
  <audio id="replyAudio" controls></audio>

  <script>
    const talkBtn = document.getElementById('talkBtn');
    const transcriptEl = document.getElementById('transcript');
    const audioEl = document.getElementById('replyAudio');

    const SpeechRecognition = window.SpeechRecognition || window.webkitSpeechRecognition;
    const recognizer = new SpeechRecognition();
    recognizer.lang = 'en-US';
    recognizer.interimResults = false;

    talkBtn.onclick = () => {
      recognizer.start();
      talkBtn.disabled = true;
      talkBtn.textContent = 'Listening...';
    };

    recognizer.onresult = async (event) => {
      const spoken = event.results[0][0].transcript;
      transcriptEl.textContent = `You said: "${spoken}"`;
      talkBtn.disabled = false;
      talkBtn.textContent = 'Start Listening';

      // TODO: replace this with your LLM call; for demo we just echo back
      const reply = `You just said: ${spoken}`;

      const resp = await fetch('http://localhost:3000/synthesize', {
        method: 'POST',
        headers: { 'Content-Type': 'application/json' },
        body: JSON.stringify({ text: reply })
      });
      const data = await resp.json();
      audioEl.src = `data:audio/mpeg;base64,${data.audio}`;
      audioEl.play();
    };

    recognizer.onerror = (e) => {
      console.error(e);
      talkBtn.disabled = false;
      talkBtn.textContent = 'Start Listening';
    };
  </script>
</body>
</html>
Enter fullscreen mode Exit fullscreen mode

Open index.html in a browser, click Start Listening, and speak. The assistant will repeat what you said using ElevenLabs‑generated speech.

Voice Cloning in a Few Minutes

One of ElevenLabs’ standout features is the ability to create a custom voice from as little as 30 seconds of audio. Here’s a quick rundown:

  1. Record a clean sample – Use a decent microphone and record a short monologue (e.g., a paragraph from a favorite book).
  2. Upload via the dashboard – Navigate to the “Voice Lab” section and click “Create new voice”. Paste the affiliate link https://try.elevenlabs.io/kr07zfuqn1bp to get a free trial credit if you haven’t already.
  3. Grab the new voice ID – After processing, ElevenLabs returns a UUID. Replace the VOICE_ID constant in server.js with this new ID.

Now your assistant will speak with your voice, your colleague’s voice, or even a fictional character’s voice—all without the need for a deep learning pipeline.

Adding Real Conversational Intelligence

The demo above simply echoes the user’s input. In a production app you’ll want a language model to generate meaningful replies. Here’s a minimal example using OpenAI’s gpt-3.5-turbo:

// add to server.js (install openai: npm i openai)
const { OpenAI } = require('openai');
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

async function getChatResponse(message) {
  const completion = await openai.chat.completions.create({
    model: 'gpt-3.5-turbo',
    messages: [{ role: 'user', content: message }],
  });
  return completion.choices[0].message.content.trim();
}

// Inside /synthesize route, replace the static reply:
const reply = await getChatResponse(text);
Enter fullscreen mode Exit fullscreen mode

Now the flow becomes:

  1. User speaks → STT → text.
  2. Server sends text to LLM → gets answer.
  3. Answer is fed to ElevenLabs → audio → back to the browser.

Debugging Tips

Issue Likely Cause Fix
401 Unauthorized from ElevenLabs Wrong or missing API key Verify ELEVENLABS_API_KEY in .env and that the key is active.
No audio returned, empty base64 string Accept: audio/mpeg missing or responseType not set Ensure responseType: 'arraybuffer' and Accept header are present.
Voice sounds robotic Using default voice with low stability Increase stability (0.7‑0.9) and similarity_boost.
Long latency (>3 s) Large text payload or network throttling Split long paragraphs into smaller chunks and synthesize sequentially.

Deploying to the Cloud

When you’re ready to go public, you can push the Express app to services like Vercel, Render, or Railway. Because the TTS endpoint streams binary data, make sure the platform supports responseType: 'arraybuffer'. You’ll also need to set environment variables (ELEVENLABS_API_KEY, OPENAI_API_KEY, etc.) in the dashboard of your chosen host.

Wrap‑Up

Building an AI voice assistant is now a matter of stitching together a few well‑documented APIs:

  • STT – Web Speech API (or any cloud provider)
  • LLM – OpenAI, Cohere, Anthropic, etc.
  • TTS – ElevenLabs for natural speech and voice cloning

The heavy lifting—high‑quality synthesis and voice cloning—is handled by ElevenLabs, letting you ship a polished product in days instead of weeks.

Ready to give your assistant a human voice? Try ElevenLabs today via this link and get a free credit to start experimenting: https://try.elevenlabs.io/kr07zfuqn1bp

Top comments (0)