DEV Community

Cover image for Streaming YouTube into a Discord voice channel: the four bugs that nearly killed it
Vihanga
Vihanga

Posted on

Streaming YouTube into a Discord voice channel: the four bugs that nearly killed it

Shen web app

Shen is a self-hosted music bot I run for my own server. You search a song on the web app, hit play, and it either streams in your browser or lands in a Discord voice channel — same track, same queue, one process.

Live at shen.zagan.space. Source is private for now.

This isn't a tutorial. It's the four problems that took real time, and the fixes.

The shape of it

One Node process, one repo, two deployment targets:

server.js      Express: REST API + Server-Sent Events + serves the built site + boots the bot
bot.js         discord.js + @discordjs/voice: joins channels, streams audio, 14 slash commands
audio.js       yt-dlp → ffmpeg → the browser player, and YouTube captions → synced lyrics
queue.js       per-guild queues + the web "pending" buffer (both are EventEmitters)
search.js      YouTube + YouTube Music search, tolerant of layouts changing under it
store.js       JSON persistence: pending, library, history, playlists, favorites
ytdlp-util.js  the single place that decides how yt-dlp talks to YouTube
web/           Vite + React 19 + Tailwind v4 + shadcn/ui
Enter fullscreen mode Exit fullscreen mode

Backend on Render (Docker, because it needs ffmpeg and yt-dlp binaries), frontend on Vercel with /api/* rewritten to the backend. The container entrypoint is literally:

CMD ["sh", "-c", "wireproxy -c /app/warp/wireproxy.conf & sleep 3 && node server.js"]
Enter fullscreen mode Exit fullscreen mode

You'll see why that exists in a minute.

Live state is a plain SSE stream. Every mutation emits change on the queue manager, one function serializes a snapshot, every connected browser gets it:

queueManager.on("change", broadcast);

app.get("/api/events", (req, res) => {
  res.set({
    "Content-Type": "text/event-stream",
    "Cache-Control": "no-cache",
    Connection: "keep-alive",
  });
  // 15s comment-ping so proxies don't reap an idle stream
  const heartbeat = setInterval(() => res.write(": ping\n\n"), 15000);
  ...
});
Enter fullscreen mode Exit fullscreen mode

No WebSocket library, no shared mutable state on the client, and the queue UI in the browser is always exactly what the bot thinks the queue is.

1. "Sign in to confirm you're not a bot"

This is the wall. yt-dlp worked perfectly from my laptop and returned nothing at all from Render — not a 403, just zero formats. YouTube blocks datacenter IP ranges at the playability check, before it even looks at cookies.

Two changes fixed it. First, don't claim to be a browser that isn't, and try several player clients:

export async function ytdlpBaseArgs() {
  const args = [
    "--no-warnings",
    "--no-progress",
    "--force-ipv4",
    "--extractor-args", "youtube:player_client=web,mweb,android",
    "--user-agent", UA,
  ];
  if (await isWarpAvailable()) args.push("--proxy", WARP_PROXY);
  return args;
}
Enter fullscreen mode Exit fullscreen mode

Second — and this is the part that actually mattered — change the exit IP. I run a Cloudflare WARP userspace proxy (wireproxy, a single static binary) inside the container and point yt-dlp at socks5://127.0.0.1:1080. The availability check is a 2-second TCP connect that gets cached, so when the proxy isn't running the app degrades to a direct connection instead of hanging:

export async function isWarpAvailable() {
  if (warpAvailable !== null) return warpAvailable;
  warpAvailable = await new Promise((resolve) => {
    const sock = net.connect({ host, port });
    sock.setTimeout(2000, () => done(false));
    sock.once("connect", () => done(true));
    sock.once("error", () => done(false));
  });
  return warpAvailable;
}
Enter fullscreen mode Exit fullscreen mode

If you take one thing from this post: when a scraper works at home and returns nothing in production, suspect your egress IP before you suspect the scraper.

2. Silence, but with a green status light

Discord voice wanted Opus. I was handing it Opus. It played pure silence.

The reason is in the ffmpeg invocation. If you declare an already-encoded ogg/Opus stream as StreamType.Opus, the decoder assumptions don't line up and you get zeros. Give it raw PCM instead and let @discordjs/voice run its own encoder:

const ffmpegArgs = [
  "-i", "pipe:0",
  "-f", "s16le",       // raw PCM, not a container
  "-ar", "48000",      // Discord's native sample rate
  "-ac", "2",
  "-acodec", "pcm_s16le",
  "-map_metadata", "-1",
  "-loglevel", "error",
  "pipe:1",
];

ytdlp.stdout.pipe(ffmpeg.stdin);
// only resolve once real audio is flowing, not when the process spawns
ffmpeg.stdout.once("data", () => {
  resolve({ stream: ffmpeg.stdout, proc: ffmpeg, inputType: StreamType.Raw });
});
Enter fullscreen mode Exit fullscreen mode

Note the last line. Spawning successfully means nothing — I only start the AudioPlayer after the first byte of PCM arrives, with a 30s timeout to kill both children. Otherwise a slow extraction shows up as an instant "nothing is playing".

3. One dropped connection took down the whole server

Render started returning 502s at random. The log said EPIPE.

In Node, an error event on a stream with no listener is thrown. Not handled — thrown, globally. My audio pipeline had four streams (ytdlp.stdout, ytdlp.stderr, ffmpeg.stdin, ffmpeg.stdout) and when someone closed their laptop mid-track, ffmpeg.stdin emitted EPIPE with nobody listening, and the entire bot + API process died. Restart, wait, repeat.

The fix is unglamorous and mandatory. Attach a listener to every child-process stream:

const safeOnError = (label, stream) =>
  stream.on("error", (err) => console.warn(`[bot] ${label} stream error: ${err.message}`));
safeOnError("ytdlp stdout", ytdlp.stdout);
safeOnError("ytdlp stderr", ytdlp.stderr);
safeOnError("ffmpeg stdin", ffmpeg.stdin);
safeOnError("ffmpeg stderr", ffmpeg.stderr);
safeOnError("ffmpeg stdout", ffmpeg.stdout);
Enter fullscreen mode Exit fullscreen mode

And a safety net at the top of the process, because one missed stream should never be an outage:

process.on("uncaughtException", (err) => console.error("[server] uncaughtException:", err));
process.on("unhandledRejection", (reason) => console.error("[server] unhandledRejection:", reason));
Enter fullscreen mode Exit fullscreen mode

I'd rather have a degraded bot than a 502 page. Treat every child-process stream as capable of throwing at you, because they will.

4. The bug that was actually a security hole

The web app had a "send this to Discord" button. It POSTed a guildId and channelId to my API, and the API passed them straight to the bot.

Which means any visitor could POST the ID of any server the bot was in and steer it into any voice channel. It's a single-user app, so it didn't feel like an auth problem — but it was.

The fix inverts the trust direction. The client never gets to say where. The server derives the destination from the caller's own Discord voice state, which the bot already tracks:

function requireVoiceAuth(req, res) {
  const user = sessionUser(req);
  if (!user?.id) {
    res.status(401).json({ error: "Sign in with Discord required" });
    return null;
  }
  const voice = getUserVoiceState(user.id);
  if (!voice) {
    res.status(400).json({ error: "You're not in a voice channel. Join a voice channel in Discord first." });
    return null;
  }
  // guildId + channelId come from Discord's own voice state, never the request body
  const guild = client.guilds.cache.get(voice.guildId);
  return { guildId: voice.guildId, channelId: voice.channelId, guild };
}
Enter fullscreen mode Exit fullscreen mode

You can only control the server you're actually connected to, because the server checked. Session IDs live in an in-memory Map behind an HttpOnly cookie — I originally packed the user object straight into the cookie and it grew past the 4KB limit, which silently broke every login callback. Moving session state server-side fixed it in one commit.

Two more worth mentioning

The double-add. Clicking "Queue" on the website added the track to the web's pending buffer and fired the bot's play path, so every track appeared twice. The fix is a single decision in one place:

async enqueue(track) {
  const active = this.activeQueue();
  if (active && active.playing) {
    return active.add({ ...track, queuedVia: "website" }); // live, exactly once
  }
  return addPending(track); // idle: buffer until /join or /play
}
Enter fullscreen mode Exit fullscreen mode

Live queue if the bot is playing, pending buffer if it isn't. Never both.

The double-start race. When a track ended, q.next() emitted a change event that triggered playback, and the player's own Idle handler also triggered playback. Same track, started twice, queue advancing twice. A one-line reentrancy guard fixed it:

const starting = new Set(); // guards the gap between "ended" and "next stream up"

async function playTrack(guildId, track) {
  if (starting.has(guildId)) return;
  starting.add(guildId);
  try { /* ... */ } finally { starting.delete(guildId); }
}
Enter fullscreen mode Exit fullscreen mode

The part nobody warns you about

YouTube reorganised its search results page and my search broke in production while my last commit was green. play-dl aborts the entire batch when it hits a single channel or Shorts card, so I stopped trusting it and parse ytInitialData myself, skipping anything that isn't a plain video:

const v = item?.videoRenderer;
if (!v?.videoId) continue;                       // not a video: channel, playlist, shelf
if (v.upcomingEventData?.startTime) continue;     // scheduled premiere
if (!v.lengthText?.simpleText) continue;          // live stream, no duration
Enter fullscreen mode Exit fullscreen mode

Same story for playlists, which moved to a new lockupViewModel shape, and for YouTube Music, which goes through the InnerTube WEB_REMIX endpoint. I wrapped every third-party call that ships with no timeout — play-dl's validate, playlist_info, all_videos — in my own, because one slow URL will otherwise hang a request for minutes and hold a connection open:

function withTimeout(promise, ms, label) {
  return new Promise((resolve, reject) => {
    const timer = setTimeout(() => reject(new Error(`${label} timed out after ${ms / 1000}s`)), ms);
    promise.then((v) => { clearTimeout(timer); resolve(v); },
                 (e) => { clearTimeout(timer); reject(e); });
  });
}
Enter fullscreen mode Exit fullscreen mode

Assume every scraper is one deploy away from breaking, and write the fallback parser before you need it.

One non-code bug worth writing down

Vercel builds failed with vite: command not found. The cause was my own: I'd run npm install --package-lock-only in web/ to quiet a dependency warning, and it pruned vite out of the lockfile. Vercel runs npm ci, which installs exactly what the lockfile says — so no Vite.

The trap that cost me an hour: npm run build still succeeded locally, because node_modules on my machine still had the old Vite installed. A green local build tells you nothing about your lockfile. Regenerate it, confirm node_modules/vite is in the lockfile, commit it.

What I'd do next

  • A real database. Everything is JSON files rewritten on every mutation. Fine for one user, wrong the moment there's a second.
  • Remove play-dl from bot.js. The comment claims it's a stream fallback; the code only ever uses yt-dlp. It's still load-bearing in search.js, where it is.
  • Drop googleapis and opusscript. Both are dependencies I stopped using.
  • Multi-user separation. Right now the bot has one active guild at a time. Proper per-user queues with real auth is the obvious architectural next step.

Honest limitations

  • Streaming YouTube audio to other people is against YouTube's terms and will hit rate limits at any real scale. This is a personal server tool, not a product.
  • Only the active guild's queue is surfaced on the website at a time.
  • The WARP proxy is a workaround, not a fix. YouTube could close this hole tomorrow and I'd be back to square one on a different error message.

If you're building something similar, the hard parts in order of how much time they cost me: egress IP → stream error handling → trusting the client → layout drift. Everything else was ordinary CRUD.

Top comments (0)