DEV Community

Cover image for How to Extract Tools and Technologies from YouTube Videos Automatically
WgeorgeAssistantIA
WgeorgeAssistantIA

Posted on

How to Extract Tools and Technologies from YouTube Videos Automatically

If you watch technical YouTube videos for work, you know the problem: a 40-minute walkthrough mentions a dozen tools, a few commands and a couple of versions, and an hour later you cannot remember which ones. This post shows a practical pipeline to extract the tools and technologies mentioned in a video automatically, step by step, and the pitfalls that make the naive version unreliable.

The pipeline in four steps

  1. Get the transcript
  2. Split it into chunks the model can handle
  3. Extract candidate tools with their timestamps
  4. Verify each candidate and attach an official link

Each step has a failure mode. We will go through them in order.

Step 1: get the transcript

Most videos have captions, either written by the author or generated automatically. Libraries such as yt-dlp can fetch them along with timing data. Keep the timing: every segment of text has a start time, and that is what lets you link each tool back to the exact moment of the video later.

Pitfall: automatic captions mangle product names. Words like "Supabase", "Vercel" or "Kubernetes" come out phonetically. Do not clean them yet; the model will need context to recover them.

Step 2: split the transcript into chunks

Long videos do not fit in one prompt, and even when they do, extraction quality drops as the input grows. A simple approach is to cut the transcript into chunks of fixed size, run them in parallel and merge the results.

We found chunk size matters more than expected. In our tests on a transcript of roughly 475,000 characters, small chunks found more tools than large ones: very large chunks returned fewer distinct tools, because the model summarises instead of listing. Smaller chunks cost more calls but give better recall. A middle size of about 30,000 characters gave a good balance between the number of calls and the tools found.

The mistake to avoid is silent truncation. An early version of our pipeline cut the text at a fixed length and analysed only the first minutes of a seven-hour video without telling anyone. If you cap the input, say so in the output.

Step 3: extract candidates with timestamps

Ask the model for structured output, not prose. For each candidate tool, request the name, a short description, a category and a verbatim quote from the transcript with its timestamp:

{
  "tools": [
    {
      "name": "Supabase",
      "category": "database",
      "quote": "we are going to use Supabase for the backend",
      "timestamp": "00:12:41"
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Two rules improve the results a lot. First, define what counts as a tool in the prompt, and list exclusions (generic words, channel names, company names that are not products). Second, require a verbatim quote. A candidate whose quote does not appear in the transcript is rejected.

After merging the chunks, deduplicate by normalised name and drop candidates without a valid timestamp.

Step 4: verify and link

This is the step most tutorials skip, and the one that matters most. A name alone is ambiguous: "Neon", "Serena" or "Codex" each match several unrelated things on the web. If you resolve the official website with a naive search, you will get cinema studios and food standards instead of software.

What worked for us:

  • Build the search query from the tool's category and description, not only its name
  • Accept canonical open-source repositories when a tool has no website of its own
  • Block aggregator sites, course platforms, forums and social networks as official sources
  • Score the match, and return nothing when the confidence is low

The last point is the hardest culture shift: an empty link is better than a wrong one. Users forgive a missing link; they do not forgive a confident wrong one.

Putting it together

A good extraction pipeline is mostly about discipline: keep the timestamps, never truncate silently, require quotes, and prefer no link to a bad link. Models are the easy part. The rest is verification.

If you want to see this applied rather than build it yourself, we built VidScope to explore this problem: paste a YouTube link and it returns the tools mentioned with timestamps and official links.

FAQ

Can I extract tools from any YouTube video?
From any video that has a transcript, though automatic captions can misspell product names.

Why keep timestamps?
They let you check each tool at the moment it was mentioned.

Why not just ask an AI to summarise the video?
A summary tells you the idea; it rarely gives you the exact tools, versions and the evidence for each.

Top comments (1)

Collapse
 
koda2026 profile image
Harun - solo dev •

wgeorge, this is a brilliantly pragmatic breakdown of a pipeline that so many developers overcomplicate.

requiring a verbatim quote before accepting a candidate is a masterclass in preventing llm hallucination. it acts as a strict grounding mechanism, ensuring the model isn't just inventing tools based on phonetic auto-caption mangling.

your finding that ~30k character chunks yield better recall than massive chunks is also a crucial insight. it perfectly mirrors the token-window limitations we face in constrained environments. forcing the model to focus on a smaller context window prevents it from defaulting to lazy summarization, which is a common trap in long-context rag.

but the absolute standout takeaway is: "an empty link is better than a wrong one." this is the exact same "fail closed" philosophy that should govern all ai systems. confidently serving a wrong link destroys user trust infinitely faster than a missing link ever could.

great work on vidscope. this level of disciplined, verification-first pipeline design is exactly what separates toy scripts from production-ready tools. 🐯