TL;DR
We'll build a dedup worker that catches re-encoded, resized, watermarked and trimmed copies of videos you already have. Stage 1 is a byte hash, stage 2 is a 64-bit perceptual hash checked by Hamming distance, stage 3 is FFmpeg's built-in MPEG-7
signaturefilter for "does A contain B, and where". Python 3.12, FFmpeg 7 or newer, SQLite.
What we're building
A dedup.py module with one entry point, check(path) -> Verdict, that returns exact, near, contains, or new, plus a SQLite index so it gets faster as the library grows. No ML, no GPUs, nothing that isn't pip install or apt install.
The reason for three stages instead of one: the cheap check should run on every upload before you spend money transcoding, and the expensive check should only run on candidates the cheap check surfaced.
1. Set up
sudo apt-get install -y ffmpeg # 7.x or 8.x; the signature filter has shipped since 2017
python3 -m venv .venv && source .venv/bin/activate
pip install videohash2 # maintained fork of videohash; pin the current release in requirements.txt
ffmpeg -hide_banner -filters | grep signature
You should see:
... signature N->V Calculate the MPEG-7 video signature
If that line is missing, your FFmpeg build was configured without it; grab a static build.
2. Make some test inputs 🎬
We'll synthesise a source and three "duplicates" so the tests are reproducible and the repo stays binary-free.
# scripts/make_fixtures.sh
set -e
mkdir -p fixtures && cd fixtures
# 20-second source
ffmpeg -y -f lavfi -i "testsrc2=duration=20:size=640x360:rate=25,format=yuv420p" \
-f lavfi -i "sine=frequency=440:duration=20" \
-c:v libx264 -crf 20 -c:a aac -shortest source.mp4
# A: re-encode at lower quality and resolution
ffmpeg -y -i source.mp4 -vf scale=320:180 -c:v libx264 -crf 32 -c:a aac dup_reencode.mp4
# B: watermark in the corner
ffmpeg -y -i source.mp4 -vf "drawbox=x=10:y=10:w=80:h=30:color=white@0.8:t=fill" \
-c:v libx264 -crf 23 -c:a copy dup_watermark.mp4
# C: 6-second excerpt starting at 0:07 (containment case)
ffmpeg -y -ss 7 -i source.mp4 -t 6 -c:v libx264 -crf 23 -c:a aac excerpt.mp4
# D: unrelated video
ffmpeg -y -f lavfi -i "mandelbrot=size=640x360:rate=25" -t 20 \
-c:v libx264 -crf 23 unrelated.mp4
sha256sum fixtures/*.mp4 gives you four different hashes for what a human would call two videos. That's the problem.
3. Stage 1: the byte hash (keep it, it's free)
# dedup/stage1.py
import hashlib
from pathlib import Path
def sha256(path: Path, chunk: int = 1 << 20) -> str:
h = hashlib.sha256()
with path.open("rb") as f:
while blk := f.read(chunk):
h.update(blk)
return h.hexdigest()
This catches the accidental double upload and nothing else. It stays because it costs nothing and it's the only stage with zero false positives.
4. Stage 2: perceptual hash + Hamming gate 🔍
videohash2 samples one frame per second, shrinks each to 144x144, tiles them into a collage, and wavelet-hashes the collage into 64 bits. Two re-encodes of the same clip land a few bits apart.
# dedup/stage2.py
from pathlib import Path
from videohash2 import VideoHash
def phash(path: Path) -> int:
vh = VideoHash(path=str(path))
return int(vh.hash_hex, 16) # 64-bit int
def hamming(a: int, b: int) -> int:
return (a ^ b).bit_count()
Try it on the fixtures:
>>> from pathlib import Path
>>> from dedup.stage2 import phash, hamming
>>> src = phash(Path("fixtures/source.mp4"))
>>> for name in ["dup_reencode", "dup_watermark", "excerpt", "unrelated"]:
... print(name, hamming(src, phash(Path(f"fixtures/{name}.mp4"))))
You'll see the re-encode and the watermark land close to zero and unrelated land far away. The interesting one is excerpt: it will not be close. That's by design and the library says so in its docs: a whole-video collage hash cannot tell you that one video is a part of another. That's what stage 3 is for.
⚠️ Note: pick the Hamming threshold conservatively and log near-misses. A false positive here means one user's upload gets served as another user's asset. Start tight (single digits), review what you're rejecting, and loosen only with evidence.
5. Stage 3: MPEG-7 signature for containment
FFmpeg's signature filter fingerprints every frame and, given two or more inputs, searches for matching sequences. First, generate and store a signature per asset. Generation is the slow part; the maintainer of one tool built on this filter reports 19 minutes to fingerprint 96 files versus 45 seconds to compare them all. Never regenerate.
# one-time per asset: write a binary signature next to it
ffmpeg -hide_banner -i fixtures/source.mp4 \
-vf "signature=format=binary:filename=fixtures/source.sig" -f null -
Then compare two inputs in one run:
ffmpeg -hide_banner -i fixtures/source.mp4 -i fixtures/excerpt.mp4 \
-filter_complex "signature=nb_inputs=2:detectmode=full" -f null - 2>&1 | grep -i match
Expected shape of the output (offsets will differ):
[Parsed_signature_0 @ 0x...] matching of video 0 at 7.000000 and 1 at 0.000000, 150 frames matching
It found the excerpt inside the source and told you where. Run the same against unrelated.mp4 and you get no matching line at all. Run it against dup_reencode.mp4 and you get whole video matching.
Wrap it:
# dedup/stage3.py
import re, subprocess
from pathlib import Path
MATCH = re.compile(r"matching of video 0 at ([\d.]+) and 1 at ([\d.]+), (\d+) frames matching")
WHOLE = re.compile(r"whole video matching")
def signature_compare(a: Path, b: Path) -> dict | None:
cmd = ["ffmpeg", "-hide_banner", "-nostats", "-i", str(a), "-i", str(b),
"-filter_complex", "signature=nb_inputs=2:detectmode=full:th_di=50",
"-f", "null", "-"]
out = subprocess.run(cmd, capture_output=True, text=True).stderr
m = MATCH.search(out)
if not m:
return None
return {
"offset_a": float(m.group(1)),
"offset_b": float(m.group(2)),
"frames": int(m.group(3)),
"whole": bool(WHOLE.search(out)),
}
th_di=50 sets the minimum matching sequence to 50 frames (two seconds at 25 fps) so a single similar frame doesn't count. The filter's other thresholds (th_d, th_dc, th_xh, th_it) have sane defaults; leave them until you have a labelled set to tune against.
6. The denominator decision
frames alone isn't a verdict. You have to divide by something, and the choice is a policy, not a constant.
| Denominator | 6s excerpt in 20s source | Answers |
|---|---|---|
| source frames (500) | 150 / 500 = 30% | "Is most of my video in theirs?" |
| shorter video's frames (150) | 150 / 150 = 100% | "Is one of these inside the other?" |
For dedup you want the second. Most hand-rolled implementations use the first because it's the natural thing to write, and then wonder why excerpts never get flagged.
# dedup/verdict.py
def coverage(frames: int, frames_a: int, frames_b: int) -> float:
return frames / min(frames_a, frames_b)
Get frame counts from ffprobe -v error -count_frames -select_streams v:0 -show_entries stream=nb_read_frames -of csv=p=0 file.mp4, or cheaper, duration * r_frame_rate from a plain probe.
7. Put it together with a SQLite index
# dedup/worker.py
import sqlite3
from dataclasses import dataclass
from pathlib import Path
from .stage1 import sha256
from .stage2 import phash, hamming
from .stage3 import signature_compare
from .verdict import coverage
PHASH_MAX_DISTANCE = 6
MIN_COVERAGE = 0.4
@dataclass
class Verdict:
kind: str # exact | near | contains | new
match_id: int | None = None
detail: dict | None = None
def check(path: Path, db: sqlite3.Connection, frames_of) -> Verdict:
digest = sha256(path)
row = db.execute("SELECT id FROM assets WHERE sha256=?", (digest,)).fetchone()
if row:
return Verdict("exact", row[0])
h = phash(path)
candidates = []
for aid, other_hash, other_path in db.execute("SELECT id, phash, path FROM assets"):
d = hamming(h, int(other_hash))
if d <= PHASH_MAX_DISTANCE:
return Verdict("near", aid, {"hamming": d})
if d <= 20: # loose band: worth a stage-3 look
candidates.append((aid, Path(other_path)))
for aid, other in candidates:
m = signature_compare(other, path)
if m and coverage(m["frames"], frames_of(other), frames_of(path)) >= MIN_COVERAGE:
return Verdict("contains", aid, m)
db.execute("INSERT INTO assets(path, sha256, phash) VALUES (?,?,?)",
(str(path), digest, str(h)))
db.commit()
return Verdict("new")
Schema:
-- dedup/schema.sql
CREATE TABLE IF NOT EXISTS assets (
id INTEGER PRIMARY KEY,
path TEXT NOT NULL,
sha256 TEXT NOT NULL UNIQUE,
phash TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_phash ON assets(phash);
💡 Tip: the linear scan over
phashis fine for tens of thousands of assets. Past that, bucket by the top 16 bits of the hash, or move to a BK-tree, before you reach for a vector database.
Run it over the fixtures in order:
python -m dedup.cli fixtures/source.mp4 fixtures/dup_reencode.mp4 \
fixtures/dup_watermark.mp4 fixtures/excerpt.mp4 fixtures/unrelated.mp4
# source.mp4 new
# dup_reencode.mp4 near (hamming=<small>)
# dup_watermark.mp4 near (hamming=<small>)
# excerpt.mp4 contains (offset_a≈7.0, frames≈150, coverage≈1.00)
# unrelated.mp4 new
# exact distances and frame counts depend on your FFmpeg build; the verdicts should not
8. Things to know
-
Store signatures as derived assets, like thumbnails. Write the
.sigonce and keep it next to the media; stage 3 becomes comparison-only. - Stage 2 fails on rotation and reversal. Rotation beyond roughly ten degrees or a reversed clip produces a different hash. If that matters for you, generate hashes for the four rotations at index time.
- Stage 3 is block-based and the FFmpeg docs say it is "not very robust to some geometric modification". Heavy crops and picture-in-picture will slip through. It's still far better than nothing.
- Sampling rate is a cost knob. One tool's experiment moving from 5 fps to 15 fps sampling saw 1.53x the signature-generation time, 5.61x the comparison time, and about 3x the storage. Start low.
-
Run in log-only mode first. Record verdicts for a week without acting on them, then look at the
nearandcontainsrows by hand before you let the worker skip an encode.
What's next
- Hook
check()into your upload webhook or queue consumer so it runs before the transcode job is enqueued; the whole point is skipping the encode. - Add a
referencetable for takedown assets and run stage 3 against it on every upload, not just nightly. - If you want to see the frames instead of the numbers,
ffmpeg -ss <offset_a> -i source.mp4 -frames:v 1 a.pngnext to the same from the candidate is a five-second visual sanity check.
Tag it #python and #ffmpeg if you write up your own thresholds; the tuning data is the part nobody shares.
Top comments (0)