Most AI video summaries read well and can't be checked. A summary of an 18-minute lecture says "the network uses a sigmoid", and you have no way of knowing whether that came from minute 10, from minute 17, or from nowhere. This post walks through a different approach on a real lecture, plus a 15-line script for searching a transcript export.
The problem with summaries that float free of the source
A summary has one job: let you skip most of the video and still know what it says. That only works if you can trust it, and trust is the part most summarizers skip. They hand back a tidy paragraph with no way to see which sentence in the video produced which claim. When something looks off, the only way to check is to scrub through the video yourself, which is exactly the work the summary was supposed to save.
For entertainment that is fine. For a lecture you are studying, a conference talk you want to quote, or a tutorial you are about to follow step by step, it is not. A wrong claim costs you more than a missing one, because you act on it.
The fix I settled on is simple to state: every point in the summary carries the timestamp it came from. The summary stops being an answer you have to believe and becomes an index into the video. You read it in a minute, and anything that matters, you check at the source in one click.
A timestamp-linked summary is a summary in which every point carries the moment in the video where it was said, so each claim can be checked against the original in one click instead of being taken on trust.
The test case: 3Blue1Brown's neural network lecture
To make this concrete I'll use a video most developers have either watched or been told to watch: But what is a neural network? | Deep learning chapter 1 by 3Blue1Brown. It runs 18:40, it's dense, and it is the kind of video people bookmark and never finish.
Here is the shortest version, the Gist, as summarizer.video produced it:
This video breaks down the mathematical architecture of a basic multilayer neural network designed to classify handwritten digits. It details how visual data flows from pixel activations through hidden layers via weighted sums, biases, and activation functions, explaining the mechanics using matrix algebra and contrasting classical sigmoid functions with modern ReLU activations.
That is enough to decide whether the video is worth your evening. If it is, the next layer down is where timestamps start doing the work.
Three depths, and every point has a time
The same summary comes at three lengths, made together from one transcript: Gist (a few sentences), Brief (the argument) and Full notes (careful notes, as if you watched with a pen). Gist is for deciding. Brief is for remembering. Full notes are for when you need the details but not the 18 minutes.
In Brief and Full notes, each point sits next to its timestamp:
Four points from the Full notes, word for word, with the times they link to:
| Time | Point |
|---|---|
| 02:17 | A plain vanilla neural network is explored to build foundational understanding for how neural networks are structured and learn without relying on buzzwords. |
| 03:04 | Neurons in the network simply hold a numerical activation value between 0 and 1, with 784 neurons in the input layer corresponding to grayscale pixels of a 28x28 image. |
| 03:47 | The final layer consists of 10 neurons representing the digits 0 through 9, while intermediate hidden layers process and propagate these activations. |
| 07:15 | A layered architecture ideally enables hierarchical feature detection, where lower layers identify small edges, intermediate layers combine edges into loops or lines, and final layers assemble digits. |
Clicking a point plays the video from that second and lights up the matching transcript line, so "784 neurons" takes a couple of seconds to verify. The transcript line at 03:08 reads: "each of the 28x28 pixels of the input image, which is 784 neurons in total."
Chapters turn a long video into a table of contents
For an 18-minute video, chapters matter more than the summary. They tell you where things are, which is what you need on the second viewing:
Twelve sections, from "Introduction example" at 00:00 through "Edge detection example" at 08:38, "Counting weights and biases" at 11:34 and "Notation and linear algebra" at 13:26, to "ReLU vs Sigmoid" at 17:03. If all you want is the matrix notation, that's 13:26 to 15:17, and you can skip the rest without guilt.
Getting the transcript out, and a script to search it
The transcript sits beside the summary and exports in five formats: text with timestamps, plain text, Markdown, and SRT or VTT subtitles with real cue timings.
SRT is the useful one for scripting, because every cue has a start time. I keep this little function around for the question I ask of almost every talk: where exactly did they say X?
import sys
def find(srt_path, term):
"""Print the start time and text of every cue that mentions `term`."""
blocks = open(srt_path, encoding="utf-8").read().strip().split("\n\n")
for block in blocks:
lines = block.splitlines()
if len(lines) < 3:
continue
start = lines[1].split(" --> ")[0].split(",")[0] # 00:03:08,120 -> 00:03:08
text = " ".join(lines[2:])
if term.lower() in text.lower():
print(f"{start} {text}")
if __name__ == "__main__":
find(sys.argv[1], sys.argv[2])
Run against the SRT for this lecture (286 cues), it answers "when does sigmoid come up?":
$ python find_in_srt.py neural-network.srt sigmoid
00:10:32 And a common function that does this is called the sigmoid function,
00:11:15 weighted sum before plugging it through the sigmoid squishification function.
00:11:54 on to the weighted sum before squishing it with the sigmoid.
00:14:43 Then as a final step, I'll wrap a sigmoid around the outside here,
00:14:50 sigmoid function to each specific component of the resulting vector inside.
00:15:55 and which involves iterating many matrix vector products and the sigmoid
00:17:15 So Lisha one thing I think we should quickly bring up is this sigmoid function.
00:17:30 Exactly. But relatively few modern networks actually use sigmoid anymore.
00:18:11 Using sigmoids didn't help training or it was very difficult to
That output is a reading plan: the definition at 10:32, the matrix form at 14:43, and the "few modern networks use this anymore" turn at 17:30. Nine lines instead of eighteen minutes.
Mind map and flashcards for the second pass
Two more views are built from the same transcript, and both keep the timestamps. The mind map lays the lecture out as topics (Multi-Layer Architecture at 02:53, Hierarchy of Abstraction at 05:49, Weights and Biases at 09:09, Linear Algebra Formulation at 13:41) and downloads as a PNG:
Flashcards are the part I use most for anything I actually want to remember. This lecture produced ten, each tied to a moment in the video, and they export as a tab-separated file that Anki and Quizlet import directly:
"What role does a bias play in computing a neuron's activation?" links to 11:21. If you get it wrong, the answer is one click away, in the lecturer's own words.
What it does not do
A few honest limits, so you know when to reach for something else:
- Private and age-restricted videos can't be opened. If you have the file, upload it instead (mp4, mov, mp3 or wav, up to 100 MB).
- The summary is only as good as the transcript. Captions are used as they are; videos without captions go through speech-to-text, which is most accurate on clear audio. For YouTube videos without captions, that needs a paid plan.
- Free accounts take videos up to 10 minutes. Opening a video that is already on the site is free and needs no account, which is how you can open the lecture above.
Try it on the video you keep putting off
I built summarizer.video around one rule: a summary should be quicker than the video and still let you check every point at its source. It works with YouTube, TikTok and Instagram links and your own files. If there's a talk sitting in your Watch Later that you will never get to, paste it into the AI video summarizer and read it first. Then watch only the part you need.






Top comments (0)