If you build anything with LLMs on top of video — summaries, search, a RAG chatbot over a course — the first problem is getting clean transcripts. Most tools give you one big block of text. For RAG you want pieces of ~250 words that still know when they were said, so your bot can cite the exact second.
Here is the shortest path I found, using the YouTube Transcript API Actor on Apify (no YouTube API key, no cookies):
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("om_kh/youtube-transcript-api").call(run_input={
"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"outputFormats": ["chunks", "srt"],
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
for chunk in item["chunks"]:
print(chunk["start"], chunk["end"], chunk["text"][:80])
Each chunk carries start and end in seconds, so a RAG answer can link to https://youtu.be/<id>?t=<start>. The same run also returns a ready-to-save .srt string, useful for video editors.
Other formats
-
vttfor web players - plain text for summaries
- whole channels or playlists: the YouTube Channel Transcripts Actor takes an
@handle
From an AI agent (MCP)
It also works as an MCP tool through https://mcp.apify.com, so Claude, Cursor or any MCP client can call it directly: "get the transcript of this video and summarize it".
Pricing
Pay per transcript delivered (see the Actor page for the current price); videos without captions are never charged.
More step-by-step API guides (Google Trends, YouTube comments, job boards, Kalshi & Polymarket odds): github.com/omarkhandji-commits/apify-actor-catalogue
Disclosure: I built this Actor. Happy to help if you hit issues — leave a comment.
Top comments (0)