DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Using ElevenLabs in Electron Desktop Apps

Why add voice to your Electron app?

Desktop apps have been dominated by clicks, forms, and visual cues for years. Adding a natural‑sounding voice layer can turn a static UI into an assistant that guides users, reads notifications, or even narrates long‑form content. With the rise of AI‑powered text‑to‑speech (TTS) and voice cloning, it’s now feasible to ship a personalized voice experience that feels human, not robotic.

If you’ve been looking for a TTS service that offers high‑quality neural voices, low latency, and a straightforward HTTP API, ElevenLabs is a solid choice. Their models produce expressive speech and even let you create custom voice clones – perfect for branding or accessibility features.

👉 Grab a free trial and start cloning voices here: https://try.elevenlabs.io/kr07zfuqn1bp

In this article we’ll walk through wiring ElevenLabs into an Electron desktop app, from project scaffolding to streaming audio playback. By the end you’ll have a minimal but functional “Speak‑It‑Out” demo that you can extend for your own product.


Prerequisites

What you need Why
Node.js ≥ 18 Modern JavaScript features and native fetch support
Electron ≥ 25 Stable APIs for the main/renderer processes
ElevenLabs API key Authenticates your requests – get one from the link above
Basic HTML/CSS/JS The UI will be a simple button and text area

If you haven’t installed Electron before, the quick start looks like this:

# Create a folder and initialize npm
mkdir elevenlabs-electron && cd $_
npm init -y

# Install Electron as a dev dependency
npm i -D electron@latest
Enter fullscreen mode Exit fullscreen mode

Create a main.js (the main process) and an index.html (the renderer). We’ll fill them out next.


Setting up the Electron skeleton

main.js

// main.js – Electron's main process
const { app, BrowserWindow, ipcMain } = require('electron');
const path = require('path');

function createWindow() {
  const win = new BrowserWindow({
    width: 500,
    height: 400,
    webPreferences: {
      preload: path.join(__dirname, 'preload.js'), // secure bridge
    },
  });

  win.loadFile('index.html');
}

// When Electron is ready, create the window
app.whenReady().then(createWindow);

// Graceful shutdown on macOS
app.on('window-all-closed', () => {
  if (process.platform !== 'darwin') app.quit();
});
Enter fullscreen mode Exit fullscreen mode

preload.js

// preload.js – expose a safe API to the renderer
const { contextBridge, ipcRenderer } = require('electron');

contextBridge.exposeInMainWorld('electronAPI', {
  speak: (text) => ipcRenderer.invoke('speak', text),
});
Enter fullscreen mode Exit fullscreen mode

index.html

<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8" />
  <title>ElevenLabs Voice Demo</title>
  <style>
    body { font-family: sans-serif; padding: 20px; }
    textarea { width: 100%; height: 120px; }
    button { margin-top: 10px; padding: 10px 20px; }
  </style>
</head>
<body>
  <h2>Speak It Out</h2>
  <textarea id="txt" placeholder="Type something..."></textarea>
  <br />
  <button id="speakBtn">🔊 Speak</button>

  <script src="renderer.js"></script>
</body>
</html>
Enter fullscreen mode Exit fullscreen mode

renderer.js

// renderer.js – runs in the renderer process
const txt = document.getElementById('txt');
const btn = document.getElementById('speakBtn');

btn.addEventListener('click', async () => {
  const text = txt.value.trim();
  if (!text) return alert('Please enter some text.');

  // Ask the main process to generate and play audio
  await window.electronAPI.speak(text);
});
Enter fullscreen mode Exit fullscreen mode

At this point you can run npx electron . and you’ll see a tiny window with a textarea and a button. Nothing happens yet because we haven’t wired the TTS call.


Calling ElevenLabs from the main process

The heavy lifting lives in main.js. We’ll use the native fetch API (available in recent Node versions) to POST the text to ElevenLabs, receive an MP3 stream, and pipe it to the OS audio output using the play-sound npm package.

npm i play-sound
Enter fullscreen mode Exit fullscreen mode

Now extend main.js:

// Add these imports near the top
const fetch = require('node-fetch'); // Node < 18 needs this, otherwise skip
const player = require('play-sound')({});

// Your ElevenLabs API key – keep it secret!
const ELEVEN_API_KEY = process.env.ELEVEN_API_KEY; // set via .env or launch script

// Helper to build the request body
function buildPayload(text) {
  return {
    text,
    voice_settings: {
      stability: 0.75,
      similarity_boost: 0.85,
    },
  };
}

// IPC handler for "speak"
ipcMain.handle('speak', async (event, text) => {
  const url = 'https://api.elevenlabs.io/v1/text-to-speech/EXAMPLE_VOICE_ID/stream';
  // Replace EXAMPLE_VOICE_ID with the ID of the voice you want, e.g., "eleven_monolingual_v1"
  const response = await fetch(url, {
    method: 'POST',
    headers: {
      'xi-api-key': ELEVEN_API_KEY,
      'Content-Type': 'application/json',
      Accept: 'audio/mpeg',
    },
    body: JSON.stringify(buildPayload(text)),
  });

  if (!response.ok) {
    const err = await response.text();
    console.error('ElevenLabs error:', err);
    throw new Error('Failed to synthesize speech');
  }

  // Stream the MP3 directly to a temporary file
  const { writeFile, unlink } = require('fs').promises;
  const tmpPath = `${app.getPath('temp')}/speech-${Date.now()}.mp3`;
  const buffer = await response.buffer();
  await writeFile(tmpPath, buffer);

  // Play the file using the OS default player
  return new Promise((resolve, reject) => {
    player.play(tmpPath, (err) => {
      // Clean up the temp file after playback
      unlink(tmpPath).catch(() => {});
      if (err) reject(err);
      else resolve();
    });
  });
});
Enter fullscreen mode Exit fullscreen mode

A few notes

  1. Voice ID – ElevenLabs ships with a few ready‑made voices (eleven_monolingual_v1, eleven_multilingual_v1). You can also create a custom clone via their dashboard. Paste the ID in the URL where EXAMPLE_VOICE_ID lives.
  2. Environment variables – Never hard‑code the API key. Store it in a .env file and load it with dotenv (npm i dotenv). Then add require('dotenv').config(); at the top of main.js.
  3. Streaming vs. whole‑file – The endpoint we used (/stream) returns an MP3 stream. For large texts you might want to pipe the response directly to the player to avoid buffering the whole file in memory. The demo keeps it simple with a temporary file.

Handling longer utterances and chunking

ElevenLabs imposes a maximum character limit per request (usually ~5 000 characters). If you need to read articles or documentation, split the text into manageable chunks:

function chunkText(text, max = 4000) {
  const chunks = [];
  while (text.length > max) {
    // Find a space near the limit to avoid cutting words
    const cut = text.lastIndexOf(' ', max);
    const part = text.slice(0, cut);
    chunks.push(part);
    text = text.slice(cut).trim();
  }
  if (text) chunks.push(text);
  return chunks;
}
Enter fullscreen mode Exit fullscreen mode

You can then loop through each chunk, awaiting the speak IPC call sequentially. This gives a smooth, continuous narration without hitting API limits.


Security & best practices

Concern Recommendation
API key exposure Keep the key in the main process only. Never expose it to the renderer or bundle it into the packaged app.
Rate limiting ElevenLabs enforces per‑minute quotas. Cache repeated utterances or debounce UI actions.
User privacy If you let users upload custom voice clones, store the clone IDs securely and provide an opt‑out.
Packaging Use electron-builder or electron-forge to create distributables. Ensure the temporary audio files are written to a safe location (app.getPath('temp')).

Going further

  • Voice cloning: Upload a few seconds of a speaker’s audio via the ElevenLabs dashboard, retrieve the new voice ID, and let users pick their “avatar” voice.
  • Realtime chat assistants: Combine the TTS flow with a LLM (e.g., OpenAI or Claude) to build a conversational desktop assistant.
  • Accessibility: Pair the audio output with screen‑reader friendly markup to serve users with visual impairments.
  • Custom UI: Replace the basic button with a floating “talk” icon, or integrate with Electron’s system tray for quick voice commands.

Wrap‑up

Adding expressive speech to an Electron app is now a matter of a few lines of code and a reliable TTS provider. By leveraging ElevenLabs’ neural models, you can deliver a polished voice experience that feels native, whether you’re building a productivity helper, an educational reader, or a game narrator.

Give the demo a spin, experiment with different voice IDs, and start thinking about how voice can enhance your next desktop product.

Ready to bring natural‑sounding AI voices to your app?

Grab your API key and try ElevenLabs today: https://try.elevenlabs.io/kr07zfuqn1bp

Top comments (0)