Getting the transcript of a YouTube video is easy on your laptop. Run the same code on a server, a notebook in the cloud or a scheduled job, and after a few requests YouTube answers with:
Sign in to confirm you're not a bot
That's YouTube blocking data-center IPs. I hit it myself after a handful of requests from a cloud machine while building this. The usual fix is rotating residential proxies, which is a small project on its own.
In this post I'll get transcripts from Python with that part handled on the server side, in formats that are ready for LLMs.
Disclosure: I built the tool used here, a YouTube Transcript Scraper on Apify. It costs $0.003 per transcript (residential proxies included); videos without captions are not charged. Apify's free plan includes monthly credits.
Setup
pip install apify-client
export APIFY_TOKEN=your_token
Transcripts as Markdown with timestamps
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("fguiraud/youtube-transcript-scraper").call(run_input={
"videos": [
"https://youtu.be/arj7oStGLkU",
"https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"https://youtube.com/shorts/U5-dE2nCzjM",
],
"languages": ["en"],
"outputs": ["markdown", "segments"],
})
for v in client.dataset(run.default_dataset_id).iterate_items():
if v["status"] != "ok":
print(f"[{v['status']}] {v['url']}")
continue
kind = "auto" if v["isAutoGenerated"] else "human"
print(f"{v['title']} ({v['language']}, {kind} captions, {v['wordCount']} words)")
Real output (September 2026):
[no-captions] https://www.youtube.com/watch?v=U5-dE2nCzjM
Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster) (en, human captions, 487 words)
Inside the Mind of a Master Procrastinator | Tim Urban | TED (en, human captions, 2277 words)
The Short has no captions (common for Shorts and music without speech), so it's reported and not charged.
The Markdown puts a timestamp at the start of each paragraph:
# Inside the Mind of a Master Procrastinator | Tim Urban | TED
**[00:00:12]** So in college, I was a government major, which means I had to write a lot of papers. ...
That matters for LLMs: when you ask for a summary, the model can cite where in the video each point is.
Summarize with an LLM
prompt = f"Summarize this talk in 5 bullet points, citing timestamps:\n\n{v['markdown']}"
# send `prompt` to the model of your choice
For long videos and RAG, ask for "chunks": ~1,000-character pieces with start and end times and a token estimate, ready to embed.
Other formats and languages
-
"outputs": ["srt"]or["vtt"]gives subtitle files for video editors. -
"languages": ["es", "en"]returns Spanish captions when they exist and English otherwise (esalso matches regional tracks likees-419). - Human captions are preferred over auto-generated ones;
availableLanguageslists every track of the video.
What it doesn't do
It reads the captions YouTube already has. It doesn't translate into languages the video has no captions for, and it can't read private, members-only or age-restricted videos (those come back as unavailable, not charged). For your own audio or video files without captions, a Whisper-based transcriber is the right tool.
Scripts (Markdown files, SRT, bulk lists from a .txt) are in the data-tools repository. Questions or feature ideas? Comments are open.
Top comments (0)