If you've tried fetching YouTube transcripts from a server, you've probably seen it: the same code that works on your laptop fails on AWS, GCP or Vercel. YouTube asks data-centre IPs to "sign in to confirm you're not a bot", and its web player now needs a proof-of-origin token for captions.
I needed transcripts for a RAG pipeline, so I built a small hosted scraper that handles this: YouTube Transcript Scraper on Apify. It reads captions through YouTube's app endpoint and retries on residential IPs when a server IP is challenged.
Three lines of Python
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("automationnation/youtube-transcript-scraper").call(run_input={
"videos": ["https://www.youtube.com/watch?v=UF8uR6Z6KLc", "https://www.youtube.com/@3blue1brown"],
"maxVideosPerChannel": 5,
})
docs = list(client.dataset(run["defaultDatasetId"]).iterate_items())
Each row has:
-
transcript: the full text -
segments:[{"start": 6.05, "duration": 1.0, "text": "..."}]for citing timestamps -
title,channel,viewCount,durationSeconds,description,keywords -
transcriptLanguage,isAutoGeneratedand every caption language the video has -
status:ok,no_transcriptorunavailable(onlyokrows are charged)
Into a vector store
chunks = []
for d in docs:
if d["status"] != "ok":
continue
window, start = [], d["segments"][0]["start"]
for s in d["segments"]:
window.append(s["text"])
if len(" ".join(window)) > 800:
chunks.append({"text": " ".join(window), "url": f'{d["url"]}&t={int(start)}s', "title": d["title"]})
window, start = [], s["start"]
Every chunk links back to the exact second of the video, so your answers can cite timestamps.
Cost
$1.50 per 1,000 transcripts, and videos without captions are free. Apify's free tier ($5 a month) covers about 3,300 transcripts.
From an AI agent
Add https://mcp.apify.com/?tools=automationnation/youtube-transcript-scraper to Claude or Cursor and ask "Summarise the last five videos on @lexfridman with timestamps."
Top comments (0)