DEV Community

quiethand098
quiethand098

Posted on

Chunking YouTube transcripts for RAG: timestamps, citations and missing captions

Most RAG demos over video content start with "get the transcript", and that step breaks more often than the embedding step. Three things matter: getting timestamps, chunking so every chunk can be cited, and not paying for videos with no captions.

1. Keep timestamps through chunking

If you split a transcript by characters you lose where each piece came from. Keep the caption segments (text, start, duration) and build chunks by walking segments until you reach your size limit, ending on a sentence boundary. Store the start of the first segment with the chunk.

def chunk(segments, max_chars=1000):
    out, cur, start = [], "", None
    for s in segments:
        if start is None:
            start = s["start"]
        cur += " " + s["text"].strip()
        if len(cur) >= max_chars and cur.rstrip()[-1:] in ".?!":
            out.append({"start": start, "text": cur.strip()})
            cur, start = "", None
    if cur.strip():
        out.append({"start": start, "text": cur.strip()})
    return out
Enter fullscreen mode Exit fullscreen mode

2. Make citations clickable

Store a deep link with each chunk: https://www.youtube.com/watch?v=VIDEO_ID&t=SECONDS. When your assistant answers, it can link to the exact moment in the video instead of the whole hour-long talk.

3. Handle missing captions without failing the batch

Some videos have no captions or have them disabled. Collect those IDs in an errors list and move on; do not retry them forever. Prefer the requested language track, then auto-generated, then any track.

Skipping the plumbing

I packaged all of the above as an Apify actor: paste video URLs, IDs or a playlist, set chunkChars, and get text, timestamped segments and sentence-aligned chunks with deep links. No API key or cookies, and you pay only for transcripts actually delivered.

YouTube Transcript Extractor on Apify

from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("quiethand098/youtube-transcript-rag-extractor").call(run_input={
    "videos": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
    "chunkChars": 1000,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    for c in item["chunks"]:
        print(c["link"], c["text"][:80])
Enter fullscreen mode Exit fullscreen mode

Source and notes: https://github.com/quiethand098/youtube-transcript-extractor

Top comments (2)

Collapse
 
reidmarlow profile image
Reid Marlow •

Adding a hard length ceiling inside that loop saves you when auto-generated captions skip punctuation altogether. Without an absolute cutoff, videos where Whisper or YouTube omit periods will refuse to split until the buffer hits the end of the entire transcript. Appending a chunk whenever the segment gap exceeds two seconds also catches natural speaker transitions that sentence detectors miss.

Collapse
 
quiethand098 profile image
quiethand098 •

Good points, thanks. The actor already has a hard ceiling (a chunk is closed at 1.5x the target size even without a sentence end), exactly for punctuation-free auto captions. I just shipped your second idea too: a pause of 2+ seconds between segments now closes the chunk once it is at least half full, so speaker/topic switches land on chunk boundaries more often.