<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: enrico gep</title>
    <description>The latest articles on DEV Community by enrico gep (@enrico_gep_b145b3149b358a).</description>
    <link>https://dev.to/enrico_gep_b145b3149b358a</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172167%2F8cc9c8b9-4946-42d4-8895-8201132518ad.jpg</url>
      <title>DEV Community: enrico gep</title>
      <link>https://dev.to/enrico_gep_b145b3149b358a</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/enrico_gep_b145b3149b358a"/>
    <language>en</language>
    <item>
      <title>From a 345-video YouTube playlist to a news digest grouped by topic</title>
      <dc:creator>enrico gep</dc:creator>
      <pubDate>Thu, 08 Oct 2026 22:48:12 +0000</pubDate>
      <link>https://dev.to/enrico_gep_b145b3149b358a/from-a-345-video-youtube-playlist-to-a-news-digest-grouped-by-topic-84o</link>
      <guid>https://dev.to/enrico_gep_b145b3149b358a/from-a-345-video-youtube-playlist-to-a-news-digest-grouped-by-topic-84o</guid>
      <description>&lt;p&gt;I follow a lot of YouTube channels about AI, tech and current affairs, and I make daily news podcasts in Italian. Watching everything is impossible, so I built a small pipeline that does the first pass for me: it takes a whole playlist, summarizes every video and groups the summaries by topic into one HTML page I can read in ten minutes.&lt;/p&gt;

&lt;p&gt;This is what the output looks like (a sample: 8 of the 43 topics from yesterday's run, only the AI and tech ones):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq72dyjkkh9plirlydbzz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq72dyjkkh9plirlydbzz.png" alt="Topic digest generated from a YouTube playlist" width="800" height="667"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each topic has a short synthesis written by the LLM and the list of videos that talk about it, linked to YouTube.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Transcripts with Supadata.&lt;/strong&gt; The playlist IDs and the native captions come from the &lt;a href="https://supadata.ai/r/SD768VWX" rel="noopener noreferrer"&gt;Supadata&lt;/a&gt; API through its Python SDK (&lt;code&gt;client.youtube.playlist.videos()&lt;/code&gt; and &lt;code&gt;client.transcript(..., mode="native", lang=...)&lt;/code&gt;). At this stage no video or audio is downloaded. The language is inferred from the title and checked against the &lt;code&gt;lang&lt;/code&gt; field of the response, because some videos have dubbed tracks in other languages. Requests go through a small rate limiter (the free plan allows 1 request per second) and every transcript is cached on disk, so a restart only fetches what is missing. On this playlist, fetching a transcript took about 4 seconds per video (median).&lt;/p&gt;

&lt;p&gt;I wrote a separate tutorial with the full, working code for this step: PASTE_TUTORIAL_LINK_HERE&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Fallback for videos without captions.&lt;/strong&gt; Only those get their audio downloaded with &lt;code&gt;yt-dlp&lt;/code&gt; and transcribed locally with &lt;code&gt;faster-whisper&lt;/code&gt; (&lt;code&gt;large-v3-turbo&lt;/code&gt;, batched, on a 6 GB RTX 3060: about 13 seconds for 12 minutes of audio).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Summaries.&lt;/strong&gt; Each transcript is summarized in Italian by an LLM (GLM 5.3 Flash in my setup), four videos at a time. Long transcripts are split into 6,500-word chunks. The videos I care most about get an extra pass that checks the summary against the transcript for missing points.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Grouping by topic.&lt;/strong&gt; The summaries are grouped in batches of 20, then the batch topics are merged into the final list. This merge was the weak spot: with 345 videos the single merge call got too big and timed out twice, so for this run I did the final merge with Claude instead. Next version: save each batch result to disk and merge in two levels.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. HTML page.&lt;/strong&gt; A tiny script turns the Markdown digest into one self-contained HTML page with light and dark themes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Numbers from yesterday's run
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;345 videos in the playlist, 343 summarized (2 had no usable text);&lt;/li&gt;
&lt;li&gt;43 topics, from Claude Code "mods" and new models to science and history;&lt;/li&gt;
&lt;li&gt;transcripts were the fast part; the LLM summaries are now the bottleneck, and that's the next thing I'm optimizing.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What I'd tell anyone building something similar
&lt;/h2&gt;

&lt;p&gt;Get the text from captions first and keep speech-to-text as a fallback: it's faster, cheaper and avoids downloading hundreds of files. Always check the language of what you get back. And cache every intermediate result, because a long batch can always get interrupted.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: both Supadata links in this post (above and below) are referral links. If you sign up through either one, I receive free credits.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Supadata: &lt;a href="https://supadata.ai/r/SD768VWX" rel="noopener noreferrer"&gt;https://supadata.ai/r/SD768VWX&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbh2morsb9e4ba2p5ag92.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbh2morsb9e4ba2p5ag92.png" alt=" " width="800" height="880"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>ai</category>
      <category>showdev</category>
      <category>youtube</category>
    </item>
  </channel>
</rss>
