DEV Community

Cover image for How I Built a Reverse Video Search Engine on One Home PC
Nebulous
Nebulous

Posted on

How I Built a Reverse Video Search Engine on One Home PC

Someone posts a ten second clip with no title. The first reply is always the same: "source?" Most of the time, nobody answers.

You'd think reverse image search solved this years ago. It didn't, and the reason is kind of interesting. Google Lens, TinEye and Yandex index pictures that live on web pages. A frame from the middle of a video was never on a web page, so there's nothing for them to find. The majority of big players also filter explicit results on purpose.

So I built a reverse video search of my own. It's called SauceMatch, and it finds the source of adult clips. That's about the hardest version of this problem, because the big engines filter that content out by design. The site is 18+ due to the nature of possible results. This article focuses on how i made the site, the techincal work done in the background, and a LOT of testing.

It runs on one PC at home. As I write this, the index holds about 32 million frames from a bit over 4 million videos. Growing everyday.

Here's what this post covers:

  • fingerprints that survive crops, overlays and recompression
  • 32 million vectors searched straight from disk, because I ran out of RAM
  • a geometric check that turns "looks similar" into "this is the same video"
  • moving the whole search into the visitor's browser, matcher included, compiled to WebAssembly

A video is just a pile of frames

The whole design falls out of one observation: a video is a set of pictures. If you can search pictures, you can search videos.

  • An indexed video is a handful of stills, about 20 at most.
  • A query clip is a set of frames picked from it. A screenshot is a single frame.
  • Every query frame looks up its nearest stored frames, and the hits are grouped by video.

That last step matters more than it looks. The same video gets uploaded to five sites, and all five copies come back as one result with five links. Most search engines treat duplicates as noise and throw them away. Here they're the answer. People want the original, and if the first link is dead, they want the next one.

Where the frames come from

Video sites publish preview images for every video: the thumbnail in the listing, and the frames that play when you hover over it. Those are ordinary images, and that's the only reason an index like this is possible at all.

Each frame is stored small (512 px on the longest side, WebP) and deduplicated by a hash of its decoded pixels. The same frame from three reposts is stored once and linked to three videos.

One rule saved a lot of grief: a frame that shows up in more than five videos is ignored at query time. Those are intro cards, logos and placeholder images. Without that rule, every studio intro matches every other video from the same studio.

Copy detection, not "similar images"

The obvious first move is to grab a general purpose embedding model like CLIP. It's the wrong tool. CLIP encodes what's in a picture. Ask it about a frame and it will happily hand you a thousand frames of similar rooms with similar people in them. I didn't want "similar". I wanted "the same picture, after somebody edited it."

That's a different problem, and it has a name: image copy detection. I use SSCD, a copy detection model from Meta built on ResNet-50. It turns a frame into 512 numbers, and it's trained so that a picture and its edited copies land close together: crops, scaling, color shifts, text overlays, recompression. The vectors are L2 normalized and compared by cosine similarity.

It's very good at this. It isn't perfect. The same room filmed on a different day can still land close. So a fingerprint hit is never treated as an answer on its own. More on that below.

32 million vectors on a PC under my desk

The index is FAISS IVF-PQ:

  • 16,384 inverted lists, 1,536 of them probed per query frame
  • 2,000 candidates per frame, reranked with 8 bit scalar quantized copies of the vectors (512 bytes each)
  • about 0.6 KB per vector per build, extended with the new frames every 12 hours

One detail worth knowing: an HNSW coarse quantizer lost 15% recall with inner product in my tests, so the quantizer is flat.

The RAM wall

Then the index outgrew the machine. Adding RAM wasn't an option, so the server stopped loading the index and started memory mapping it instead: the inverted lists through FAISS's IO_FLAG_MMAP_IFC, and the 8 bit codes and ids through numpy maps.

The server's private memory went from 18.1 GiB to 5.3 GiB. The OS gets to treat the index as file cache and drop pages when something else needs the memory.

Great, except cold searches fell off a cliff.

One system call took a search from 18 seconds to 1

Here's what was going on. Reranking reads 2,000 rows per query frame, scattered across a multi gigabyte file. Each row is 512 bytes. Each one was its own page fault, and each fault pulled 64 to 86 KiB off the disk. A 30 frame query took 18.3 seconds.

The fix is almost embarrassingly simple. After the IVF step you already know exactly which rows you're about to read. So ask the OS for all of them in one call, before touching any of them. On Windows that's PrefetchVirtualMemory. On Linux, madvise with MADV_WILLNEED does the same job.

import ctypes
from ctypes import wintypes
import numpy as np

PAGE = 4096
k32 = ctypes.WinDLL("kernel32", use_last_error=True)
k32.GetCurrentProcess.restype = wintypes.HANDLE
k32.PrefetchVirtualMemory.argtypes = [wintypes.HANDLE, ctypes.c_size_t, ctypes.c_void_p, wintypes.ULONG]
k32.PrefetchVirtualMemory.restype = wintypes.BOOL

def prefetch_rows(arr, rows):
    """Ask Windows for every page these rows of a mapped array touch, in one call."""
    width = np.uint64(arr.strides[0])
    start = np.uint64(arr.ctypes.data) + rows.astype(np.uint64) * width
    mask = ~np.uint64(PAGE - 1)
    pages = np.unique(np.concatenate([start & mask, (start + width - np.uint64(1)) & mask]))
    entries = np.empty((len(pages), 2), np.uint64)  # WIN32_MEMORY_RANGE_ENTRY: address, size
    entries[:, 0], entries[:, 1] = pages, PAGE
    if not k32.PrefetchVirtualMemory(k32.GetCurrentProcess(), len(pages), entries.ctypes.data, 0):
        raise ctypes.WinError(ctypes.get_last_error())
Enter fullscreen mode Exit fullscreen mode

That's trimmed from the real code, which also checks bounds and logs a failed call instead of failing the search.

Cold searches went from 18.3 s to 1.1 s. Warm ones take 0.65 s. The results are identical, because the search didn't change at all. Only the order in which the disk gets asked did.

Latency on disk

One more Windows gotcha: the index builder deletes old builds while the server may still have them mapped. A plain np.load(mmap_mode="r") holds the file in a way that makes the delete fail. Opening the maps yourself with FILE_SHARE_DELETE fixes it.

Fingerprints find lookalikes. Geometry proves it.

This is the part I care about most, because it's where most tools in this space cut corners.

Every promising candidate gets checked against the query frame with plain old computer vision in OpenCV:

  1. Find SIFT keypoints in both frames, up to 4,000 each.
  2. Pair them up and keep the pairs that pass Lowe's ratio test.
  3. Fit a homography with MAGSAC. If enough pairs agree on one transform that maps your frame onto the stored one, it's the same picture, even if it was cropped, scaled or had text slapped on it.

A simplified version:

sift = cv2.SIFT_create(nfeatures=4000)
kq, dq = sift.detectAndCompute(query_gray, None)
kc, dc = sift.detectAndCompute(candidate_gray, None)

pairs = cv2.BFMatcher(cv2.NORM_L2).knnMatch(dq, dc, k=2)
good = [m for m, n in (p for p in pairs if len(p) == 2)
        if m.distance < RATIO * n.distance]

src = np.float32([kq[m.queryIdx].pt for m in good])
dst = np.float32([kc[m.trainIdx].pt for m in good])
H, mask = cv2.findHomography(src, dst, cv2.USAC_MAGSAC, max_px_error)
inliers = int(mask.sum()) if mask is not None else 0
Enter fullscreen mode Exit fullscreen mode

The 0 to 100 confidence combines the inlier count, the inlier ratio, how widely the inliers spread across the frame, and a few pixel level checks on the overlapping region (a perceptual hash, histogram correlation, SSIM). The thresholds came from held out tests. Real reposts of the same video never scored below 28 agreeing keypoints, while different videos shot on the same set stayed under the gates.

So there are two kinds of results, and the UI never mixes them up:

  • Match: the geometry check passed. This is the same video.
  • Similar: it looks close but has no geometric proof, and it's labeled as unverified.

When nothing passes, the page says nothing was found. A lot of sites in this niche show you a grid of blurred "results" and charge you to unblur them. I'd rather show an honest empty page. The long version, with pictures, is on the page about how the Match check works.

keypoints in the query frame paired with keypoints in the stored frame, all agreeing on one transform

Moving the whole search into the browser

At first the server did everything. Then I moved the search into the visitor's browser, for two reasons. The visitor's file never has to leave their device. And a home PC has better things to do than run a neural network for every visitor.

Now the browser picks the frames, fingerprints them, sends only the fingerprints to the API, and runs the geometry check itself on thumbnails it loads from a CDN.

Quantization nearly broke it

SSCD runs in ONNX Runtime Web inside a web worker, on WebGPU when the browser has it and on WASM when it doesn't.

const session = await ort.InferenceSession.create(MODEL_URL, {
  executionProviders: ['webgpu', 'wasm'],
})
const out = await session.run({ [session.inputNames[0]]: frameTensor })
Enter fullscreen mode Exit fullscreen mode

The fp32 model is a lot to download for a web page, so I quantized it. Here's how that went:

  • Full int8: cosine 0.78 against the server's fingerprint on a dark frame. Useless.
  • Static QDQ quantization: 0.69. Worse.
  • A plain fp16 graph: NaN on WebGPU.
  • Weight only int8, with the first conv, layer1 and all biases kept in fp32: 25.5 MB, and cosine 0.9977 or better against the server.

The lesson: test a quantized model on your worst inputs, not your average ones. It looked fine until it saw a dark frame.

Threads need cross origin isolation

ONNX Runtime's WASM backend only gets threads when the page is cross origin isolated, because threads need SharedArrayBuffer. That takes two headers:

Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: credentialless
Enter fullscreen mode Exit fullscreen mode

They took a frame from about 420 ms to about 140 ms. I went with credentialless rather than require-corp, and Cloudflare's Turnstile bot check still works under it.

The same matcher, compiled to WebAssembly

The geometry check has to give the same answers in the browser as on the server. So instead of rewriting it in JavaScript, I compiled the same OpenCV 4.14 code paths to WebAssembly with Emscripten: one C++ file, run in a pool of workers.

The perceptual hashes were the fiddly part. To match the Python side bit for bit, the port reproduces Pillow's grayscale conversion and its fixed point LANCZOS resize, scipy's unnormalized DCT and numpy's median. A parity script runs both matchers on the same pairs and compares the answers. They agree, except that float differences flip a SIFT descriptor rounding now and then, and very rarely a whole keypoint.

Phones and weaker devices are the exception. They send the chosen frames to the server, which runs the same steps. Every device falls back to that path if the browser path fails.

Reading a video without a video element

Picking frames from a clip sounds easy: seek a <video> element, draw it to a canvas, repeat. On phones it turned out to be the slowest part of the whole search, 59% of it, because every seek decodes forward from the previous key frame.

So the page stopped using <video> for reading. It hands the file to a worker that demuxes it with mp4box.js and decodes runs of frames with WebCodecs.

The catch: the new reader has to pick exactly the frames the old one would have, or the search results change. That meant reproducing which frame Chrome's <video> actually shows at a given time: edit lists, composition offsets, and Chrome's rule of "the last frame at or before t". In headless Chrome, both readers now produce the same positions and the same JPEG bytes. The new reader only runs in Chromium browsers, on files it understands. Anything unusual goes back to <video>, so a weird file never fails a search.

Then comes the frame picking itself:

  • Sites cut their previews at even steps through the video (9, 15, 16, 17, 20 or 60 steps, depending on the site). The scan visits those kinds of positions plus an even sweep in between, up to 160 positions whatever the length.
  • At each position it measures brightness, contrast and sharpness on a small copy. Black frames, fades and motion blur are dropped, and near duplicates are merged.
  • Black bars, and the blurred copy of the video that phone apps put around a horizontal clip, are found across many frames and cropped away.
  • It keeps the best frames, with at least one per stretch of the timeline.

Screenshots are a mess

People don't upload clean frames. They upload screenshots of an app, with a status bar, player buttons and a comment section, and the actual video sitting in a box in the middle.

So frames that are only an app screen get detected and skipped. When a video is playing inside an app, that region gets cut out and searched on its own.

Two more tricks recover a lot of near misses:

  • When a video comes up as Similar, the page samples the clip again a few seconds either side of the frame that found it. The index only holds about 20 stills per video, so the exact one may sit between two frames of the first pass.
  • When nothing matches at all, the best frames get searched again, mirrored. Reposts are flipped surprisingly often.

The boring infrastructure

  • The pages are static, on Cloudflare Pages.
  • The API runs on the home PC behind a Cloudflare Tunnel, so no port is open.
  • Thumbnails live in Bunny Storage, behind Cloudflare's cache.
  • A GPU in the same PC fingerprints new frames as they come in.

The GPU driver resets every now and then, and I still don't know why. When the CUDA context dies, the server exits with code 75 and a small supervisor starts a fresh one within seconds. The index is rebuilt every 12 hours, and the old build keeps serving until the new one is ready.

How this differs from the reverse image search you already know

Google Lens, TinEye, Yandex "Finder" sites with an unlock button SauceMatch
What it searches Pictures on web pages Often the same big engines, repackaged Its own index of video frames
Proof None Blurred results A geometry check, scored 0 to 100
When nothing is found Lookalikes Blurred "maybes" An honest empty result
Cost Free Pay to unlock Free, no account
Your file Uploaded Uploaded Stays on your device
Reposts Mixed in Unclear Every copy, grouped under one video

It's the same pipeline whether you start with a clip, a screenshot or a GIF. It works as a reverse video search, a picture search or a GIF search.

Removals, and the other way around

Anything can be removed on request, a single frame or a whole video. The same search also works the other way around: creators can use it to check where their own video was posted, and then file takedowns.

The rules are short: adults only, lawful use, and no looking up private people.

What still doesn't work

  • Coverage. It can only find what's in the index. In the first ten days of October, 12% of searches ended with a verified Match. Most misses are videos that simply aren't indexed yet.
  • Same set, same watermark. In my tests, about 1 to 2% of Matches are wrong in one specific way: two different videos shot in the same room with the same overlay. The geometry really is consistent there. The fix I have in mind, and haven't built yet, is a photometric check after the homography.
  • Paid sites. Sites that don't publish previews in public are hard to cover at all.

What I'd tell myself at the start

  • Use a copy detection model, not a general embedding. "Similar" is the wrong question.
  • Never let a nearest neighbor be the final answer. A cheap geometric check is what makes the results trustworthy.
  • When the data stops fitting in RAM, look at your access pattern before you look at hardware. One prefetch call fixed what looked like a hardware problem.
  • Test quantized models on your worst inputs.
  • If two implementations have to agree, compile the same code twice instead of writing it twice.

If you want to try it, it's at saucematch.com. It's 18+, free, and there's no account.

There's another huge part of the site, face detection/search. Maybe another article someday. :)

If anyone is reading this, im curious if anyone has run FAISS from mmap on Linux at this scale? I'm curious whether MADV_WILLNEED holds up as well as PrefetchVirtualMemory did here.

And if anyone has recommendations or suggestions on possible ways to improve the backend, im more than happy to listen :)

Top comments (0)