Why a Text‑to‑Speech Chrome Extension?
If you spend a lot of time reading articles, documentation, or long forum threads, a quick way to listen instead of scrolling can boost productivity dramatically. With the rise of high‑quality voice AI, building a custom Text‑to‑Speech (TTS) Chrome extension is easier than ever. In this guide we’ll walk through a complete, production‑ready extension that sends selected text to the ElevenLabs API, receives a natural‑sounding audio file, and plays it back instantly.
TL;DR – By the end of this post you’ll have a Chrome extension that:
- grabs highlighted text,
- calls the ElevenLabs TTS service,
- streams the audio back to the browser,
- and works on Manifest V3 (the latest Chrome extension format).
What You’ll Need
| Item | Reason |
|---|---|
| Node.js (v14+) | To run a tiny local server for testing the extension’s background script. |
| A Chrome browser (or Edge) | To load the unpacked extension. |
| ElevenLabs API key | The service that turns text into lifelike speech. Get one for free at https://try.elevenlabs.io/kr07zfuqn1bp. |
| Basic JavaScript/HTML/CSS | The extension is pure front‑end code, no heavy frameworks needed. |
Pro tip: Store your API key in a
.envfile and load it withdotenvduring development. Never commit the key to source control.
1. Set Up the Extension Skeleton
Create a new folder called tts-extension and add the following files.
manifest.json
{
"manifest_version": 3,
"name": "ElevenLabs Text‑to‑Speech",
"description": "Select text and hear it spoken instantly using ElevenLabs AI.",
"version": "1.0.0",
"permissions": [
"activeTab",
"scripting",
"storage"
],
"action": {
"default_popup": "popup.html",
"default_icon": {
"16": "icons/icon16.png",
"48": "icons/icon48.png",
"128": "icons/icon128.png"
}
},
"background": {
"service_worker": "background.js"
},
"icons": {
"16": "icons/icon16.png",
"48": "icons/icon48.png",
"128": "icons/icon128.png"
}
}
Manifest V3 uses a service worker (
background.js) instead of a persistent background page, which keeps memory usage low.
popup.html
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Read Aloud</title>
<style>
body { font-family: Arial, sans-serif; width: 250px; padding: 10px; }
button { width: 100%; padding: 8px; margin-top: 5px; }
</style>
</head>
<body>
<h4>ElevenLabs TTS</h4>
<textarea id="text" rows="4" placeholder="Select text on the page"></textarea>
<button id="speak">🔊 Speak</button>
<p id="status"></p>
<script src="popup.js"></script>
</body>
</html>
popup.js
// Grab the highlighted text from the current tab
async function getSelection() {
const [{ result }] = await chrome.scripting.executeScript({
target: { tabId: (await chrome.tabs.query({ active: true, currentWindow: true }))[0].id },
func: () => window.getSelection().toString()
});
return result;
}
document.getElementById('speak').addEventListener('click', async () => {
const statusEl = document.getElementById('status');
const textArea = document.getElementById('text');
const text = textArea.value.trim() || await getSelection();
if (!text) {
statusEl.textContent = '⚠️ No text to read';
return;
}
statusEl.textContent = '⏳ Generating audio...';
// Send message to background service worker
chrome.runtime.sendMessage({ action: 'speak', text }, (response) => {
if (response.success) {
statusEl.textContent = '✅ Playing…';
} else {
statusEl.textContent = `❌ ${response.error}`;
}
});
});
background.js
let audio = null;
// Load API key from storage (set via options page or dev console)
async function getApiKey() {
return new Promise((resolve) => {
chrome.storage.sync.get(['elevenApiKey'], (items) => {
resolve(items.elevenApiKey);
});
});
}
// Call ElevenLabs TTS endpoint
async function fetchAudio(text, voice = 'Bella') {
const apiKey = await getApiKey();
if (!apiKey) throw new Error('ElevenLabs API key not set');
const response = await fetch('https://api.elevenlabs.io/v1/text-to-speech/' + voice, {
method: 'POST',
headers: {
'Accept': 'audio/mpeg',
'Content-Type': 'application/json',
'xi-api-key': apiKey
},
body: JSON.stringify({
text,
voice_settings: { stability: 0.75, similarity_boost: 0.85 }
})
});
if (!response.ok) {
const err = await response.text();
throw new Error(`ElevenLabs error: ${err}`);
}
return response.arrayBuffer(); // raw audio bytes
}
// Play audio using the Web Audio API
function playAudio(buffer) {
const blob = new Blob([buffer], { type: 'audio/mpeg' });
const url = URL.createObjectURL(blob);
audio?.pause();
audio = new Audio(url);
audio.play();
}
// Listen for messages from popup
chrome.runtime.onMessage.addListener((msg, sender, sendResponse) => {
if (msg.action === 'speak') {
fetchAudio(msg.text)
.then(playAudio)
.then(() => sendResponse({ success: true }))
.catch(err => sendResponse({ success: false, error: err.message }));
// Keep the message channel open for async response
return true;
}
});
2. Storing the ElevenLabs API Key
You can set the key manually via the Chrome extension's Storage panel, but a quick way during development is to run this snippet in the background console:
chrome.storage.sync.set({ elevenApiKey: 'YOUR_ELEVENLABS_API_KEY' });
Replace YOUR_ELEVENLABS_API_KEY with the key you obtain from the affiliate sign‑up page at https://try.elevenlabs.io/kr07zfuqn1bp. The key is saved securely in Chrome’s sync storage, so it’s available across your devices.
3. Testing the Extension
- Open chrome://extensions → enable Developer mode.
- Click Load unpacked and select the
tts-extensionfolder. - Pin the extension icon, open any web page, highlight a paragraph, and click the extension’s popup “🔊 Speak” button.
You should hear the selected text spoken in a natural voice. If you get an error, open the extension’s background console (via Inspect views) and look for messages like “ElevenLabs error: …”.
4. A Quick cURL Example (Optional)
If you want to experiment with the ElevenLabs API outside the extension, here’s a minimal cURL call:
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/Emma" \
-H "Content-Type: application/json" \
-H "xi-api-key: YOUR_ELEVENLABS_API_KEY" \
-d '{
"text": "Hello, developer! This is ElevenLabs speaking.",
"voice_settings": { "stability": 0.7, "similarity_boost": 0.9 }
}' \
--output hello.mp3
Replace YOUR_ELEVENLABS_API_KEY with the same key you used in the extension. The command saves an MP3 file (hello.mp3) that you can play locally.
5. Going Further
-
Voice selection – Add a dropdown in
popup.htmlto let users pick from ElevenLabs’ catalog (Bella,Emma,Arnold, …). Pass the chosen voice tofetchAudio. - Chunking – For very long articles, split the text into 500‑character chunks to stay under the API’s request size limits.
-
Caching – Store generated audio in
chrome.storage.localkeyed by a hash of the text. Subsequent reads become instantaneous. - Options page – Build a small settings UI where users can paste their API key and set a default voice.
6. Publishing (Optional)
When you’re happy with the extension:
- Create a zip of the folder (
tts-extension.zip). - Head to the Chrome Web Store Developer Dashboard and upload the zip.
- Fill in the store listing, add screenshots, and submit for review.
Make sure you comply with ElevenLabs’ usage policy—especially regarding commercial redistribution of generated audio.
Wrap‑Up
You now have a fully functional Chrome extension that leverages ElevenLabs’ state‑of‑the‑art voice synthesis. The codebase is small enough to understand at a glance, yet extensible for advanced use‑cases like multi‑language support or custom voice cloning.
Ready to give your browser a voice? Grab your free ElevenLabs API key at https://try.elevenlabs.io/kr07zfuqn1bp and start building richer, more accessible web experiences today!
Top comments (0)