DEV Community

Cover image for The Magic of Prediction in Video Codecs
Elecard
Elecard

Posted on

The Magic of Prediction in Video Codecs

An easy-to-understand guide to prediction in video coding

When you watch a video in HD, you probably don't think about how much data would have to be transmitted if there were no compression. The numbers would be staggering. Yet the video reaches you quickly and in good quality — and for that, we have video codecs and one of their key tools, prediction, to thank.

Let's take a look at how it works.

Video Is Redundant — and That's Good News

Every video contains a huge amount of redundant information. Neighboring pixels within a frame are often very similar: the sky stays blue across many pixels, a wall keeps roughly the same color, and a person's face doesn't change dramatically from one pixel to the next.

But that's not all. Neighboring frames are often very similar, too. If a person is speaking on screen, the background behind them barely changes from one frame to the next. An arm moves smoothly rather than jumping randomly from one position to another.

Smart algorithms can take advantage of all this redundancy. And this is where prediction comes in.

Prediction is a process that tries to estimate what a block being encoded looks like before it is compressed. The estimate is based on data that has already been encoded, so the decoder on the other end can perform the same process and reconstruct the image.

There are two main types of prediction: intra prediction and inter prediction.

Intra Prediction: Extending the Picture

Imagine that you're encoding a small square block of an image. Some of the pixels to its left and above have already been encoded.

The idea behind intra prediction is simple: take these neighboring pixels and extend their patterns into the block being encoded at a particular angle — almost as if you were drawing the lines with a ruler.

The resulting prediction block is subtracted from the original block. If the prediction is accurate, the residual will be very small — and it is this residual that is then transformed and processed. The more accurate the prediction, the less data needs to be processed.

This tool appeared in the AVC standard, which offered seven angular prediction directions. In the next standard, HEVC, the number of angular directions increased to 33. In the latest standard, VVC, it doubled compared with HEVC. More prediction modes give the encoder more options for finding the best direction for a particular block.

Inter Prediction: Looking for the Block in the Past

Inter prediction works differently. The idea is to find the current block in frames that have already been encoded — previous frames or even future frames that have already been processed in coding order.

If a person in the frame has moved slightly to the right, the codec can find the corresponding content in a previous frame and essentially say: "This block is the same as that one, just shifted a certain number of pixels to the right."

This idea dates back to MPEG-1 and has been continuously refined ever since.

When Motion Gets Complicated: Affine Motion Compensation in VVC

A simple shift is just one type of motion. In real-world video, objects can rotate, scale, and become distorted. Classical codecs before VVC could describe only translational motion — as if a block simply slid from one position to another.

VVC introduced affine motion compensation. The block being encoded is divided into smaller 4×4-pixel sub-blocks, and each sub-block is assigned its own motion vector. This vector is not explicitly transmitted in the bitstream; instead, it is derived from a small set of parameters that are sent to the decoder.

As a result, the codec can model rotation, scaling, and translation more accurately — although with limitations: it still cannot handle very large rotation angles.

Why Action and Sports Videos Require More Bitrate Than Webinars

The reason is fairly straightforward. The simpler and more predictable the motion in a frame, the more accurately prediction can work, and the less residual data needs to be transmitted.

In a scene with a person speaking against a static background, inter prediction works extremely well: the background hardly changes, and there is very little motion. The codec can easily find the corresponding block in a reference frame, and a static background usually requires very little new information to be encoded.

A fireworks display or a fast-paced sports scene is a different story. Motion can be complex and widespread across the entire frame. Prediction is more likely to be inaccurate, residuals become larger, and more data needs to be transmitted.

That's why sports broadcasts typically require a higher bitrate than ordinary webinars at the same resolution.


Prediction is one of the key principles that allows video codecs to compress huge amounts of data while maintaining acceptable image quality. From simply finding a similar block in a neighboring frame to more sophisticated motion models in VVC, the goal remains the same: don't transmit the entire picture from scratch — transmit only what has actually changed.

The better a codec can identify and describe those changes, the more efficiently it can use every bit.

Top comments (0)