The YouTube Data API gives you titles, views and tags — but not what was said. The captions endpoint needs the channel owner's OAuth token, so for anyone else's video there is no official way to get the transcript as data.
If you're building RAG over video libraries, summarising a competitor's channel, or feeding an AI agent with podcasts, you need transcripts in bulk, with timecodes, in a format a database accepts. This post shows one way to get that.
What you get per video
The YouTube Transcript Scraper on Apify reads the caption tracks YouTube serves to any visitor — no login, no API key, no cookies — and returns one row per video:
-
transcript— the full text, pluswordCountandlanguage -
isAutoGenerated— so you know whether it's human captions or ASR - optional
segmentswith start time and duration for every line - optional ready-made SRT or WebVTT
- optional retrieval
chunkssplit by chapter, each with a start timecode — ready for embedding - video metadata on the same row: title, channel, publish date, duration, views, likes, tags, chapters
Any address works
A single video, a youtu.be link, a Short, a bare video ID, a playlist, or a channel handle like @zdfheute (capped at the number of latest videos you choose).
curl -X POST "https://api.apify.com/v2/acts/lergassy~youtube-transcript-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"videos":["https://www.youtube.com/@zdfheute"],"maxVideosPerChannel":20,"chunkForRag":true}'
The problems it was built around
Reading the issue trackers of popular transcript tools, the same complaints repeat: "No caption was found" on videos that clearly have captions, the wrong language coming back, and runs that silently return empty strings. So:
- captions are matched by language prefix, and
language: autotakes what the video really has (English first, then the video's own language, then any other track); - one failing video never stops the run — it's retried, and if it still fails you get a row that says why in words (
no-captions,private-video,age-restricted…) plus the list of languages that do exist.
Cost
$8 per 1,000 transcripts. A video without captions is a free error row, and there is no start fee — you pay for transcripts delivered, one video one charge.
Limits
- It reads captions; it doesn't transcribe audio. A video with no caption track at all comes back as a free error row. For those, a speech-to-text tool is the fallback (see the next post).
- Members-only and private videos stay private.
Disclosure: links to Apify in this post use my referral code.
Top comments (0)