Most RAG demos over video content start with "get the transcript", and that step breaks more often than the embedding step. Three things matter: getting timestamps, chunking so every chunk can be cited, and not paying for videos with no captions.
1. Keep timestamps through chunking
If you split a transcript by characters you lose where each piece came from. Keep the caption segments (text, start, duration) and build chunks by walking segments until you reach your size limit, ending on a sentence boundary. Store the start of the first segment with the chunk.
def chunk(segments, max_chars=1000):
out, cur, start = [], "", None
for s in segments:
if start is None:
start = s["start"]
cur += " " + s["text"].strip()
if len(cur) >= max_chars and cur.rstrip()[-1:] in ".?!":
out.append({"start": start, "text": cur.strip()})
cur, start = "", None
if cur.strip():
out.append({"start": start, "text": cur.strip()})
return out
2. Make citations clickable
Store a deep link with each chunk: https://www.youtube.com/watch?v=VIDEO_ID&t=SECONDS. When your assistant answers, it can link to the exact moment in the video instead of the whole hour-long talk.
3. Handle missing captions without failing the batch
Some videos have no captions or have them disabled. Collect those IDs in an errors list and move on; do not retry them forever. Prefer the requested language track, then auto-generated, then any track.
Skipping the plumbing
I packaged all of the above as an Apify actor: paste video URLs, IDs or a playlist, set chunkChars, and get text, timestamped segments and sentence-aligned chunks with deep links. No API key or cookies, and you pay only for transcripts actually delivered.
YouTube Transcript Extractor on Apify
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("quiethand098/youtube-transcript-rag-extractor").call(run_input={
"videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
"chunkChars": 1000,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
for c in item["chunks"]:
print(c["link"], c["text"][:80])
Source and notes: https://github.com/quiethand098/youtube-transcript-extractor
Top comments (2)
Adding a hard length ceiling inside that loop saves you when auto-generated captions skip punctuation altogether. Without an absolute cutoff, videos where Whisper or YouTube omit periods will refuse to split until the buffer hits the end of the entire transcript. Appending a chunk whenever the segment gap exceeds two seconds also catches natural speaker transitions that sentence detectors miss.
Good points, thanks. The actor already has a hard ceiling (a chunk is closed at 1.5x the target size even without a sentence end), exactly for punctuation-free auto captions. I just shipped your second idea too: a pause of 2+ seconds between segments now closes the chunk once it is at least half full, so speaker/topic switches land on chunk boundaries more often.