<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Akhilesh Kancharla</title>
    <description>The latest articles on DEV Community by Akhilesh Kancharla (@akhileshkancharla).</description>
    <link>https://dev.to/akhileshkancharla</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4110197%2F43604e86-2925-4754-8efd-85460a828c32.png</url>
      <title>DEV Community: Akhilesh Kancharla</title>
      <link>https://dev.to/akhileshkancharla</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/akhileshkancharla"/>
    <language>en</language>
    <item>
      <title>What Does a Neural Network Actually Receive When You Give It an Image?</title>
      <dc:creator>Akhilesh Kancharla</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:36:51 +0000</pubDate>
      <link>https://dev.to/akhileshkancharla/what-does-a-neural-network-actually-receive-when-you-give-it-an-image-cm2</link>
      <guid>https://dev.to/akhileshkancharla/what-does-a-neural-network-actually-receive-when-you-give-it-an-image-cm2</guid>
      <description>&lt;p&gt;When we look at an image, we see objects, colors, and shapes. A neural network starts with none of those ideas. Before it can process an image, the image must be represented as numbers.&lt;/p&gt;

&lt;p&gt;So what does the model actually receive?&lt;/p&gt;

&lt;h2&gt;
  
  
  An image starts with pixels
&lt;/h2&gt;

&lt;p&gt;A color image is a grid of pixels. Each pixel usually has three values describing its red, green, and blue channels. For an 8-bit RGB image, each value ranges from 0 to 255.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5fo53mcob6ba8x7uat9r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5fo53mcob6ba8x7uat9r.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;(255, 0, 0)&lt;/code&gt; is red.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;(0, 255, 0)&lt;/code&gt; is green.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;(0, 0, 255)&lt;/code&gt; is blue.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Imagine a tiny image containing just four pixels:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Left&lt;/th&gt;
&lt;th&gt;Right&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Top&lt;/td&gt;
&lt;td&gt;Red&lt;/td&gt;
&lt;td&gt;Green&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bottom&lt;/td&gt;
&lt;td&gt;Blue&lt;/td&gt;
&lt;td&gt;White&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;We can create that image and convert it into a PyTorch tensor:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;numpy&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;PIL&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Image&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;torchvision.transforms&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;v2&lt;/span&gt;

&lt;span class="n"&gt;pixels&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;array&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
        &lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;255&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;uint8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;image&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromarray&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pixels&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;transform&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Compose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ToImage&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ToDtype&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;float32&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scale&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;tensor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;transform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tensor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="c1"&gt;# torch.Size([3, 2, 2])
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tensor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dtype&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="c1"&gt;# torch.float32
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tensor&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;    &lt;span class="c1"&gt;# tensor([1., 0., 0.])
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The top-left pixel was red: &lt;code&gt;(255, 0, 0)&lt;/code&gt;. After conversion and scaling, it is &lt;code&gt;(1.0, 0.0, 0.0)&lt;/code&gt;. Its colour has not changed; we have changed how its values are represented.&lt;/p&gt;

&lt;p&gt;The two transform steps do different jobs. &lt;code&gt;ToImage()&lt;/code&gt; converts the image into a tensor-based image object. &lt;code&gt;ToDtype(torch.float32, scale=True)&lt;/code&gt; converts its values to floating-point numbers and scales them from the usual 0–255 range to 0–1. Scaling is &lt;strong&gt;not&lt;/strong&gt; performed by &lt;code&gt;ToImage()&lt;/code&gt; alone. &lt;a href="https://docs.pytorch.org/tutorials/beginner/basics/transforms_tutorial" rel="noopener noreferrer"&gt;PyTorch explains these steps in its transforms tutorial&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is the shape &lt;code&gt;[3, 2, 2]&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkytzxxy5nwp3529uk9nd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkytzxxy5nwp3529uk9nd.png" alt=" " width="800" height="350"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The tensor has three dimensions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[channels, height, width]
[   3,       2,     2  ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The three channels hold the red, green, and blue values. Each channel is a 2 × 2 grid.&lt;/p&gt;

&lt;p&gt;This can feel backward if you are used to image arrays shaped &lt;code&gt;[height, width, channels]&lt;/code&gt;. In this PyTorch image pipeline, the channel dimension comes first. Printing &lt;code&gt;tensor.shape&lt;/code&gt; is a simple way to check what your code produced before passing it to a model.&lt;/p&gt;

&lt;p&gt;Models commonly process several images together. Adding a batch dimension changes our example’s shape from &lt;code&gt;[3, 2, 2]&lt;/code&gt; to &lt;code&gt;[1, 3, 2, 2]&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;batch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tensor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;unsqueeze&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# torch.Size([1, 3, 2, 2])
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The leading &lt;code&gt;1&lt;/code&gt; means there is one image in the batch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does the tensor tell the model what the image contains?
&lt;/h2&gt;

&lt;p&gt;No. Converting an image into a tensor gives the model numbers arranged by channel and position. It does not attach labels such as “red square,” “road,” or “cat.”&lt;/p&gt;

&lt;p&gt;A model has to learn useful patterns from training data. The tensor is the input representation that makes those computations possible; it is not an interpretation of the scene.&lt;/p&gt;

&lt;p&gt;That distinction helps when debugging computer vision code. If a model behaves strangely, check the input before changing the model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the shape what the model expects?&lt;/li&gt;
&lt;li&gt;Are the channels in the expected order?&lt;/li&gt;
&lt;li&gt;Are the values in the expected range?&lt;/li&gt;
&lt;li&gt;Did preprocessing treat training and evaluation images consistently?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A mismatch in any of these can change what the model receives, even when the original image looks perfectly normal to us.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;An image becomes a tensor by turning its pixels into an organised array of numbers. In this example, a 2 × 2 RGB image became a &lt;code&gt;[3, 2, 2]&lt;/code&gt; tensor of floating-point values between 0 and 1.&lt;/p&gt;

&lt;p&gt;The next question is more interesting: &lt;strong&gt;once the model receives those numbers, how can it begin to detect a pattern such as an edge?&lt;/strong&gt; That is where convolutions come in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;References:&lt;/strong&gt; &lt;a href="https://docs.pytorch.org/tutorials/beginner/basics/transforms_tutorial" rel="noopener noreferrer"&gt;PyTorch transforms tutorial&lt;/a&gt; · &lt;a href="https://docs.pytorch.org/tutorials/beginner/basics/tensorqs_tutorial" rel="noopener noreferrer"&gt;PyTorch tensor tutorial&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>computervision</category>
      <category>machinelearning</category>
      <category>pytorch</category>
    </item>
  </channel>
</rss>
