A broken pixel metric does not crash. It keeps returning numbers. The numbers are in range, the table looks consistent, the histogram has a sensible shape, and you write a conclusion under it.
We build a tool that reframes landscape video into vertical clips, so we measure a lot of things per frame: where the face is, where the captions sit, whether the frame is empty, who is speaking. Five of those measurements were wrong in ways that no other number could have told us. Each time, what exposed the problem was opening one frame and looking at it.
This post covers the five failures and the rule we adopted afterwards.
Background: the fit layout
Most of these failures share one cause, so the layout needs explaining first.
When a 16:9 source goes into a 9:16 output, there are two broad options. fill crops the source to a tall window. fit keeps the whole source visible: the 16:9 picture sits in the middle of the vertical frame, and the empty bands above and below are filled with a blurred, enlarged copy of the same video.
+-----------+
| blurred | enlarged, blurred copy of the video
| copy |
+-----------+
| |
| real | the actual 16:9 picture
| picture |
| |
+-----------+
| blurred |
| copy |
+-----------+
The padding is useful visually. For measurement it is a trap, because:
- It is not empty. It is the video, scaled up and blurred.
- Anything a detector can find in the real picture, it can often find in the padding too, at a larger size.
- Whole horizontal bands of the frame belong to the padding, not to the content.
Keep that in mind for the metrics below.
1. The caption band detector
What we thought it measured: the vertical position of the caption band, so we could keep subjects clear of it.
What it measured: a loose contour-based detector responds to any dense, high-contrast edge structure near the bottom of the frame. In practice that meant a broadcaster's watermark, plus the outlines of hands and plates. The "band" appeared to jump around from frame to frame because the detector was latching onto different objects.
Nothing errored. We got a position for every frame.
2. The identity histogram
What we thought it measured: whether the person in frame was the same person, by comparing colour histograms of a region.
What it measured: we sampled the region just below the detected face, assuming it would be a shirt or torso. It was frequently a hand, a napkin or a plate. The histogram changed substantially while the person stayed exactly the same, and stayed flat when a different person wore similar colours.
3. The mouth and eye motion ratio
What we thought it measured: whether someone is speaking, using landmark motion around the mouth normalised by the eyes.
def mouth_activity(landmarks_seq):
# landmarks_seq: per-frame (mouth_points, eye_points)
ratios = []
for mouth, eyes in landmarks_seq:
mouth_motion = np.ptp(mouth[:, 1]) # vertical spread
eye_scale = np.linalg.norm(eyes[0] - eyes[1])
ratios.append(mouth_motion / eye_scale)
return float(np.mean(ratios))
What it measured: jaw movement. On cooking and eating footage, the frame with the highest "speaking" score was the clearest example of someone chewing. The metric is perfectly good at detecting that a mouth is moving. It cannot tell you why.
4. The empty-frame band
What we thought it measured: whether the frame is empty (no subject), by checking how much the pixels in a vertical band deviate from a flat background.
What it measured: the band we picked fell inside the fit padding. The padding is a blurred copy of the video, and in a dark scene a blurred dark video is close to one flat colour. So dark scenes with a clearly visible subject scored as empty.
band = frame[int(0.15 * h):int(0.55 * h)]
deviation = band.reshape(-1, 3).std(axis=0).mean()
# low deviation -> "empty"
# ...in a layout where that band is blurred padding
5. Face area from a face detector
This was the clearest one, and it used a good, widely available detector (YuNet, via OpenCV).
What we thought it measured: how large the face is in the frame, as a fraction of frame area. We take the largest box per frame.
detector = cv2.FaceDetectorYN.create(model_path, "", (w, h))
_, faces = detector.detect(frame)
if faces is not None:
x, y, bw, bh = max(faces, key=lambda f: f[2] * f[3])[:4]
face_area = (bw * bh) / (w * h)
What it measured: in the fit layout, the padding contains an enlarged copy of the speaker, and the detector found that face too. Because we took the largest box per frame, the enlarged copy won. In 13 of 17 sampled frames, the selected box was in the padding. The reported face area was 14.77% against a real value of about 0.15%.
Again: valid boxes, valid confidence scores, and a believable-looking distribution.
What these have in common
None of the five raised an error. Each returned values in the expected range. And each was wrong in a way that a second metric would not have revealed, because a second metric computed on the same frames would have inherited the same mistake (the padding, the watermark, the hand).
What does reveal the problem is a single frame with the measurement drawn on it.
The rule
Before a pixel metric goes into a report or a decision, look at at least one frame with the measurement overlaid. Generate a contact sheet or a strip, open it, and look.
It does not need to be a heavy process. A helper like this is enough:
def overlay_and_save(frame, boxes, path):
out = frame.copy()
for (x, y, w, h) in boxes:
cv2.rectangle(out, (int(x), int(y)), (int(x + w), int(y + h)), (0, 255, 0), 3)
cv2.imwrite(path, out)
Then open the file. The step that matters is the last one: a person looks at it.
Checklist
- Look at the highest-scoring and lowest-scoring frames for the metric. Extremes show failures fastest; our "most talkative" frame was the chewing one.
- If the layout has padding, detect the padded frames and mask the padding, or widen the measurement region to the whole frame. Do not measure inside a band that may be blurred filler.
- Keep watermarks, broadcaster graphics and burned-in lower thirds out of the measurement region.
- Apply the same detector with the same settings to both sides of any comparison. Do not use a loose detector for coverage on one side and a strict one for position on the other. If one output has black padding and another has blurred padding, a comparison that ignores this is biased.
- Show the contact sheet to someone else. Writing it to disk is not delivering it, and a second pair of eyes catches what the author has stopped seeing.
Disclosure
I work on LunarClip, a Windows app that turns long videos into short vertical clips on your own PC. These measurements come from testing its reframing. If you want to look at it: lunarclip.com.
The rule itself is tool-agnostic. If you run a metric over video and have never looked at the frame it picked, pick the worst-scoring one today and open it.
Top comments (0)