A 40-minute conference talk. A 90-minute podcast. A 12-minute tutorial that could have been a blog post. Most of what's valuable in a video is its content — and content can be read.
I've been building GlenSum, a browser extension that summarizes long videos and articles directly in the tab. This post covers the three working approaches to media summarization in a browser, the privacy trade-off most people don't notice they're making, and what "long content" actually means for an LLM.
Three ways to summarize a YouTube video (slowest to fastest)
1. The transcript, by hand
Every video with captions has a transcript: expand the description → Show transcript → select all → paste into any AI chat.
- Cost: free
- Time: 3–5 minutes of fiddly copying, and raw transcripts are messy — no punctuation, no speaker labels, timestamps everywhere
- Best for: one-off use when you can't install anything
2. Paste the URL into an AI chat
Some assistants accept a YouTube URL directly and fetch the transcript themselves. Paste the link, ask for key points.
- Cost: often limited by your plan's usage quota
- Caveat: availability varies by model and region — and now the video passes through yet another service
- Best for: occasional videos when you already pay for an assistant
3. A browser extension, one click
Extensions built for this read the transcript from the page you already have open and generate a structured summary in a side panel. The workflow becomes: open the video → click the extension → read the summary. No copying, no tab-switching.
The extras that matter: adjustable summary length, follow-up questions about the content ("what did they say about X?"), bilingual output when the talk isn't in your first language, and Markdown/PDF export so the summary lands in your notes.
The privacy trade-off most people miss
Here's the part that motivated building this in the first place. When you paste a URL or transcript into a hosted summarizer, your reading and watching history becomes someone's dataset. Article by article, video by video, a profile of what you consume accumulates on a server you don't control. That's a real cost that never shows up on the pricing page.
The alternative is BYOK — bring your own key:
- You paste an API key from your own AI provider (OpenAI, Google, a local gateway, anything OpenAI-compatible) into the extension once.
- The extension reads the page content in your browser and sends it directly to the provider you chose.
- There is no middleman server. No account. No logs on my side, because there is no "my side" — the extension ships as static code, and the network calls go from your browser to your provider.
The trade-off is honest: you manage your own key, and free-tier users can still use built-in free models without one. But nobody in the middle ever sees what you read.
What "long" actually means for an LLM
Summarizing a 2-hour talk isn't just "paste transcript into prompt". A two-hour transcript is roughly 25,000–40,000 words — technically it may fit a modern context window, but quality degrades: the model skims, drops the middle, and blends sections together.
The approach that works is map-reduce over chunks:
- Map: split the transcript (or article) into overlapping chunks at natural boundaries — chapters, speaker turns, section headings — and summarize each chunk separately.
- Reduce: merge the chunk summaries into a final structured summary, keeping section-level detail instead of a mushy average.
This is also why page-native tools beat copy-paste: the extension already knows the page structure, so it can split at real boundaries instead of arbitrary token counts.
When a summary is the wrong tool
Summaries are ideal for inverted-pyramid content: talks, podcasts, news analysis, tutorials — where information density is low and predictable. They're poor substitutes for demonstration content: if someone is showing you how to do a three-point turn, no summary captures the steering.
A good rule: summarize first to decide whether to watch — not to avoid watching things worth watching.
Try it
GlenSum is free to install and works out of the box (built-in free models, or your own key for unlimited use):
- 🌐 Site & docs: glenkit.com/glensum
- 🧩 Edge Add-ons · Chrome Web Store
Questions about the architecture — especially around BYOK flows or chunking strategies for long transcripts — happy to answer in the comments.
Top comments (0)