, Quality first (RVM ResNet50, about 51.3 MB) and an Experimental ultimate tier that downloads about 345 MB and is for non-commercial testing only. I stuck to the small one. Everything below is the Compatible model on three free clips from Wikimedia Commons, all CC0, exported on a plain white background so the mistakes are easy to spot. The short version: on one person talking it looks fine, and each thing you add to the frame, faster movement or more people, makes it look fake in a different place.
The easy case first. "Tarun speaking 01" is a man talking at an event, 464×832, phone-like framing, with a projector screen full of text behind him. At 4 seconds his hair and shoulder line are clean, and I couldn't see the edge shimmer I was bracing for. The top of the head is where I found something: a few faint specks and a ghost of the projector lettering left above his hair. On white you have to look for it. On a busy replacement background I doubt anyone would notice. On a dark one I'm less sure.
The surprise in the same clip was everyone else. The tool doesn't cut out "the speaker", it cuts out people, so the women standing to the right got kept as foreground too, with their lower bodies turned into a semi-transparent smear. In the transparent export I made of the same clip, an office chair near one of them also came through as a grey see-through shape. If your clip has a bystander, the result has a bystander, only partly there.
Clip: Wikimedia Commons "Tarun speaking 01" (CC0).
Fast arms were next. "Juggling with swing poi" is someone swinging poi under purple light. I cut a 6-second stretch with the most motion and scaled it to 720 wide. Wherever a hand moves fast, the hand and whatever it's holding go soft and half transparent, basically motion blur turned into see-through edges. At 3 seconds a patch of the background next to his back, around a door frame, stayed in the output as if it were part of him. At 5 seconds, with both arms up, the raised hands are semi-transparent. The curly hair, oddly, held its outline better than the hands did. There was no fire in this stretch, so I can't say anything about flames.
Then the one that really fell apart: "Folkloristic dance in Naples", a street performance, 1280×720, a row of musicians and a crowd behind barriers, camera slowly panning. Where people overlap at 2 seconds, the ones in the middle turn semi-transparent and a pair of legs just isn't there. At 11 seconds the whole group of about five people on the left is gone, removed along with the street behind them. The front row on the right, lighter clothes and clear outlines, stays solid the whole time. I also tried this clip on a gradient background image I made, and at 6 seconds the gradient shows through the lower half of the people in front. The far-away crowd was removed properly, for what it's worth.

Clip: Wikimedia Commons "Folkloristic dance in Naples (tammorriata)" (CC0).
So where does it look fake? From these three: the top of the head, faintly, on a clean single-person shot. Hands and anything held in them, whenever they move fast. People standing behind or between other people, especially in dark clothes against a busy background. And anyone who isn't your subject, who gets kept whether you wanted them or not. None of that is hidden, to be fair. The page's own notes say complex occlusion and fast motion can cause edge errors, and after this I believe it.
What I'd actually use it for: one person, facing the camera, not much else going on in the frame. That covers a lot of talking-head clips. I didn't try Quality first on the crowd clip, which is the obvious next test, and I don't know yet whether the bigger model would keep that left group or just lose them more neatly. The model card, to be fair, pitches Compatible as the one for typical computers, longer videos and fast previews, not as the one for crowds.
It's at https://imging.ai/ if you want to try your own clip before trusting it with anything busy.
kk61ykro0gwx0t8xnlb.png)
Top comments (0)