DEV Community

VoiceDeveloper
VoiceDeveloper

Posted on

Build a Text-to-Speech Chrome Extension

Why a Text‑to‑Speech Chrome Extension?

If you spend a lot of time reading articles, documentation, or long forum threads, a quick way to listen instead of scrolling can boost productivity dramatically. With the rise of high‑quality voice AI, building a custom Text‑to‑Speech (TTS) Chrome extension is easier than ever. In this guide we’ll walk through a complete, production‑ready extension that sends selected text to the ElevenLabs API, receives a natural‑sounding audio file, and plays it back instantly.

TL;DR – By the end of this post you’ll have a Chrome extension that:

  • grabs highlighted text,
  • calls the ElevenLabs TTS service,
  • streams the audio back to the browser,
  • and works on Manifest V3 (the latest Chrome extension format).

What You’ll Need

Item Reason
Node.js (v14+) To run a tiny local server for testing the extension’s background script.
A Chrome browser (or Edge) To load the unpacked extension.
ElevenLabs API key The service that turns text into lifelike speech. Get one for free at https://try.elevenlabs.io/kr07zfuqn1bp.
Basic JavaScript/HTML/CSS The extension is pure front‑end code, no heavy frameworks needed.

Pro tip: Store your API key in a .env file and load it with dotenv during development. Never commit the key to source control.


1. Set Up the Extension Skeleton

Create a new folder called tts-extension and add the following files.

manifest.json

{
  "manifest_version": 3,
  "name": "ElevenLabs Text‑to‑Speech",
  "description": "Select text and hear it spoken instantly using ElevenLabs AI.",
  "version": "1.0.0",
  "permissions": [
    "activeTab",
    "scripting",
    "storage"
  ],
  "action": {
    "default_popup": "popup.html",
    "default_icon": {
      "16": "icons/icon16.png",
      "48": "icons/icon48.png",
      "128": "icons/icon128.png"
    }
  },
  "background": {
    "service_worker": "background.js"
  },
  "icons": {
    "16": "icons/icon16.png",
    "48": "icons/icon48.png",
    "128": "icons/icon128.png"
  }
}
Enter fullscreen mode Exit fullscreen mode

Manifest V3 uses a service worker (background.js) instead of a persistent background page, which keeps memory usage low.

popup.html

<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <title>Read Aloud</title>
  <style>
    body { font-family: Arial, sans-serif; width: 250px; padding: 10px; }
    button { width: 100%; padding: 8px; margin-top: 5px; }
  </style>
</head>
<body>
  <h4>ElevenLabs TTS</h4>
  <textarea id="text" rows="4" placeholder="Select text on the page"></textarea>
  <button id="speak">🔊 Speak</button>
  <p id="status"></p>

  <script src="popup.js"></script>
</body>
</html>
Enter fullscreen mode Exit fullscreen mode

popup.js

// Grab the highlighted text from the current tab
async function getSelection() {
  const [{ result }] = await chrome.scripting.executeScript({
    target: { tabId: (await chrome.tabs.query({ active: true, currentWindow: true }))[0].id },
    func: () => window.getSelection().toString()
  });
  return result;
}

document.getElementById('speak').addEventListener('click', async () => {
  const statusEl = document.getElementById('status');
  const textArea = document.getElementById('text');
  const text = textArea.value.trim() || await getSelection();

  if (!text) {
    statusEl.textContent = '⚠️ No text to read';
    return;
  }

  statusEl.textContent = '⏳ Generating audio...';

  // Send message to background service worker
  chrome.runtime.sendMessage({ action: 'speak', text }, (response) => {
    if (response.success) {
      statusEl.textContent = '✅ Playing…';
    } else {
      statusEl.textContent = `❌ ${response.error}`;
    }
  });
});
Enter fullscreen mode Exit fullscreen mode

background.js

let audio = null;

// Load API key from storage (set via options page or dev console)
async function getApiKey() {
  return new Promise((resolve) => {
    chrome.storage.sync.get(['elevenApiKey'], (items) => {
      resolve(items.elevenApiKey);
    });
  });
}

// Call ElevenLabs TTS endpoint
async function fetchAudio(text, voice = 'Bella') {
  const apiKey = await getApiKey();
  if (!apiKey) throw new Error('ElevenLabs API key not set');

  const response = await fetch('https://api.elevenlabs.io/v1/text-to-speech/' + voice, {
    method: 'POST',
    headers: {
      'Accept': 'audio/mpeg',
      'Content-Type': 'application/json',
      'xi-api-key': apiKey
    },
    body: JSON.stringify({
      text,
      voice_settings: { stability: 0.75, similarity_boost: 0.85 }
    })
  });

  if (!response.ok) {
    const err = await response.text();
    throw new Error(`ElevenLabs error: ${err}`);
  }

  return response.arrayBuffer(); // raw audio bytes
}

// Play audio using the Web Audio API
function playAudio(buffer) {
  const blob = new Blob([buffer], { type: 'audio/mpeg' });
  const url = URL.createObjectURL(blob);
  audio?.pause();
  audio = new Audio(url);
  audio.play();
}

// Listen for messages from popup
chrome.runtime.onMessage.addListener((msg, sender, sendResponse) => {
  if (msg.action === 'speak') {
    fetchAudio(msg.text)
      .then(playAudio)
      .then(() => sendResponse({ success: true }))
      .catch(err => sendResponse({ success: false, error: err.message }));
    // Keep the message channel open for async response
    return true;
  }
});
Enter fullscreen mode Exit fullscreen mode

2. Storing the ElevenLabs API Key

You can set the key manually via the Chrome extension's Storage panel, but a quick way during development is to run this snippet in the background console:

chrome.storage.sync.set({ elevenApiKey: 'YOUR_ELEVENLABS_API_KEY' });
Enter fullscreen mode Exit fullscreen mode

Replace YOUR_ELEVENLABS_API_KEY with the key you obtain from the affiliate sign‑up page at https://try.elevenlabs.io/kr07zfuqn1bp. The key is saved securely in Chrome’s sync storage, so it’s available across your devices.


3. Testing the Extension

  1. Open chrome://extensions → enable Developer mode.
  2. Click Load unpacked and select the tts-extension folder.
  3. Pin the extension icon, open any web page, highlight a paragraph, and click the extension’s popup “🔊 Speak” button.

You should hear the selected text spoken in a natural voice. If you get an error, open the extension’s background console (via Inspect views) and look for messages like “ElevenLabs error: …”.


4. A Quick cURL Example (Optional)

If you want to experiment with the ElevenLabs API outside the extension, here’s a minimal cURL call:

curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/Emma" \
  -H "Content-Type: application/json" \
  -H "xi-api-key: YOUR_ELEVENLABS_API_KEY" \
  -d '{
        "text": "Hello, developer! This is ElevenLabs speaking.",
        "voice_settings": { "stability": 0.7, "similarity_boost": 0.9 }
      }' \
  --output hello.mp3
Enter fullscreen mode Exit fullscreen mode

Replace YOUR_ELEVENLABS_API_KEY with the same key you used in the extension. The command saves an MP3 file (hello.mp3) that you can play locally.


5. Going Further

  • Voice selection – Add a dropdown in popup.html to let users pick from ElevenLabs’ catalog (Bella, Emma, Arnold, …). Pass the chosen voice to fetchAudio.
  • Chunking – For very long articles, split the text into 500‑character chunks to stay under the API’s request size limits.
  • Caching – Store generated audio in chrome.storage.local keyed by a hash of the text. Subsequent reads become instantaneous.
  • Options page – Build a small settings UI where users can paste their API key and set a default voice.

6. Publishing (Optional)

When you’re happy with the extension:

  1. Create a zip of the folder (tts-extension.zip).
  2. Head to the Chrome Web Store Developer Dashboard and upload the zip.
  3. Fill in the store listing, add screenshots, and submit for review.

Make sure you comply with ElevenLabs’ usage policy—especially regarding commercial redistribution of generated audio.


Wrap‑Up

You now have a fully functional Chrome extension that leverages ElevenLabs’ state‑of‑the‑art voice synthesis. The codebase is small enough to understand at a glance, yet extensible for advanced use‑cases like multi‑language support or custom voice cloning.

Ready to give your browser a voice? Grab your free ElevenLabs API key at https://try.elevenlabs.io/kr07zfuqn1bp and start building richer, more accessible web experiences today!

Top comments (0)