Auditing Train and Test Splits in a Video Dataset
A random video split can make evaluation look strong when the same creator, series, or footage appears on both sides. A split audit should happen before model tuning, not after a surprising score.
Check identities first
Group records by platform ID, creator ID, channel, series, and source URL. Exact IDs catch obvious duplicates. Perceptual video hashes and short audio fingerprints help find edited or reposted clips.
Choose the split unit deliberately
Split by creator or channel when style, voice, background, or repeated formats could leak. Split by source group when multiple clips come from the same collection. Document the rule so the evaluation can be reproduced.
Report what was removed
Keep a decision log with canonical record, duplicate reason, method version, and review state. Report leakage candidates separately from ordinary duplicates; they affect evaluation differently.
Thordata describes structured video metadata, captions, transcripts, engagement signals, and ready-to-use or custom delivery options. Ask for sample provenance fields before integrating a dataset: https://www.thordata.com/products/multi-platform-video-datasets?op=rhea&from=x
Top comments (0)