DEV Community

Akhilesh Kancharla
Akhilesh Kancharla

Posted on

What Does a Neural Network Actually Receive When You Give It an Image?

When we look at an image, we see objects, colors, and shapes. A neural network starts with none of those ideas. Before it can process an image, the image must be represented as numbers.

So what does the model actually receive?

An image starts with pixels

A color image is a grid of pixels. Each pixel usually has three values describing its red, green, and blue channels. For an 8-bit RGB image, each value ranges from 0 to 255.

For example:

  • (255, 0, 0) is red.
  • (0, 255, 0) is green.
  • (0, 0, 255) is blue.

Imagine a tiny image containing just four pixels:

Left Right
Top Red Green
Bottom Blue White

We can create that image and convert it into a PyTorch tensor:

import numpy as np
import torch
from PIL import Image
from torchvision.transforms import v2

pixels = np.array(
    [
        [[255, 0, 0], [0, 255, 0]],
        [[0, 0, 255], [255, 255, 255]],
    ],
    dtype=np.uint8,
)

image = Image.fromarray(pixels)

transform = v2.Compose(
    [
        v2.ToImage(),
        v2.ToDtype(torch.float32, scale=True),
    ]
)

tensor = transform(image)

print(tensor.shape)       # torch.Size([3, 2, 2])
print(tensor.dtype)       # torch.float32
print(tensor[:, 0, 0])    # tensor([1., 0., 0.])
Enter fullscreen mode Exit fullscreen mode

The top-left pixel was red: (255, 0, 0). After conversion and scaling, it is (1.0, 0.0, 0.0). Its colour has not changed; we have changed how its values are represented.

The two transform steps do different jobs. ToImage() converts the image into a tensor-based image object. ToDtype(torch.float32, scale=True) converts its values to floating-point numbers and scales them from the usual 0–255 range to 0–1. Scaling is not performed by ToImage() alone. PyTorch explains these steps in its transforms tutorial.

Why is the shape [3, 2, 2]?

The tensor has three dimensions:

[channels, height, width]
[   3,       2,     2  ]
Enter fullscreen mode Exit fullscreen mode

The three channels hold the red, green, and blue values. Each channel is a 2 × 2 grid.

This can feel backward if you are used to image arrays shaped [height, width, channels]. In this PyTorch image pipeline, the channel dimension comes first. Printing tensor.shape is a simple way to check what your code produced before passing it to a model.

Models commonly process several images together. Adding a batch dimension changes our example’s shape from [3, 2, 2] to [1, 3, 2, 2]:

batch = tensor.unsqueeze(0)
print(batch.shape)  # torch.Size([1, 3, 2, 2])
Enter fullscreen mode Exit fullscreen mode

The leading 1 means there is one image in the batch.

Does the tensor tell the model what the image contains?

No. Converting an image into a tensor gives the model numbers arranged by channel and position. It does not attach labels such as “red square,” “road,” or “cat.”

A model has to learn useful patterns from training data. The tensor is the input representation that makes those computations possible; it is not an interpretation of the scene.

That distinction helps when debugging computer vision code. If a model behaves strangely, check the input before changing the model:

  • Is the shape what the model expects?
  • Are the channels in the expected order?
  • Are the values in the expected range?
  • Did preprocessing treat training and evaluation images consistently?

A mismatch in any of these can change what the model receives, even when the original image looks perfectly normal to us.

The takeaway

An image becomes a tensor by turning its pixels into an organised array of numbers. In this example, a 2 × 2 RGB image became a [3, 2, 2] tensor of floating-point values between 0 and 1.

The next question is more interesting: once the model receives those numbers, how can it begin to detect a pattern such as an edge? That is where convolutions come in.

References: PyTorch transforms tutorial · PyTorch tensor tutorial

Top comments (0)