most autograd tutorials show you requires_grad, then .backward(), then .grad, and then they hand you a training loop, and you come out of it able to run autograd and still not able to say what it did, and i think that is backwards. so in this one you do the derivative first, by hand, on a tiny network that fits on one page, and only once you have your own number do we run pytorch and check it against .grad. the two agree to the decimal. and that is the point where autograd stops being magic, because you already did the thing it does, you just did it slower.
First, what a gradient actually is
Skip this section if you are comfortable with it, but most confusion about autograd is really confusion about what it is producing.
Imagine standing on a hillside in fog. You cannot see the bottom, but you can feel that the ground tilts. Step in the direction it tilts downward and you get lower. That is a gradient: the slope of the ground under your feet.
Now put that on a graph. The horizontal axis is one adjustable number inside the model, a parameter. The vertical axis is the loss. The curve shows how the loss changes as the parameter
changes, and somewhere on it there is a lowest point.
Take y = x². Its slope at position x is 2x. At x = 3 the slope is 6, steep and positive, so the loss climbs to the right and you should step left. At x = -2 it is -4, so you step right. At x = 0 it is 0, and you have arrived.
import torch
x = torch.tensor(3.0, requires_grad=True)
y = x ** 2
y.backward()
print(x.grad) # tensor(6.)
PyTorch produced the slope without being told the formula. That is autograd, and the rest of this article is how.
Turning on the tape
Tensors do not track gradients by default. That would be wasted work on your input data, which is never adjusted. You opt in with requires_grad=True:
x = torch.tensor(3.0, requires_grad=True)
print(x.requires_grad) # True
From this moment every operation involving x gets recorded. Think of it as a receipt. Each multiplication, addition and power you apply gets written down, along with enough information to reverse it later.
The computation graph
The receipt has a proper name, the computation graph. It is a chain of nodes where each node is an operation and each edge carries a tensor from one operation into the next.
x = torch.tensor(3.0, requires_grad=True)
y = x ** 2 # node: power
z = 2 * y + 1 # nodes: multiply, then add
Three lines of Python, four values, three operation nodes:
x (leaf) **2 ×2 +1 z
3.0 → 9.0 → 18.0 → 19.0 → 19.0
This is the forward pass. It runs left to right, computes the answer, and as a side effect builds the graph. Nothing has been differentiated yet. The graph exists purely so that something can walk it backwards.
Two things are worth noticing here. First, x is a leaf — you created it directly rather than computing it from something else. Second, every result node remembers the operation that produced it. You can see this:
print(z.grad_fn) # <AddBackward0 object at ...>
print(y.grad_fn) # <PowBackward0 object at ...>
print(x.grad_fn) # None ← leaves have no history
grad_fn is the recording. AddBackward0 is not the addition, it is the instruction for reversing the addition.
Calling .backward()
z.backward()
print(x.grad) # tensor(12.)
Check it by hand. z = 2x² + 1, so dz/dx = 4x, and at x = 3 that is 12. Correct.
But the interesting part is not that the answer is right, it is how it was produced, because PyTorch never formed the expression 4x at all. It walked the graph right to left and at each node multiplied the incoming gradient by that node's own local slope:
z = 19 +1 ×2 **2 x.grad
start slope 1 slope 2 slope 2x = 6
grad 1 → grad 1 → grad 2 → grad 12
Reading right to left: the +1 node has slope 1, so the gradient passes through unchanged. The ×2 node has slope 2, so the gradient doubles. The x² node has slope 2x, which at x = 3 is 6, so the gradient multiplies by 6. Altogether 1 × 2 × 6 = 12.
That is the whole of backpropagation. It is the chain rule, applied one node at a time, right to left, automatically. Nobody ever writes down 4x.
This also explains why the forward pass has to happen first, and why it has to store its intermediate values. The x² node's local slope is 2x, which needs the value of x that went in. Every node keeps whatever it needs to compute its own derivative later. This is where the memory goes during training, and it is why a batch that fits in memory for inference can still run you out of memory when training.
Now do it on an actual network, by hand
Everything above is one variable. Here is the smallest thing that is honestly a neural network:
one input, one weight, one bias, a ReLU, and a squared-error loss. Five operations. Enough to show every part of the chain rule, small enough to hold in your head.
Numbers: x = 2.0, w = 3.0, b = -1.0, target y = 2.0.
Forward pass, left to right, writing down everything:
h = w·x + b = 3·2 + (-1) = 5.0
a = ReLU(h) = max(0, 5) = 5.0
L = (a - y)² = (5 - 2)² = 9.0
Loss is 9. Now the question that matters: how should w and b change to make it smaller?
Backward pass, one node at a time, right to left. At each node, local derivative times the
gradient arriving from the right.
Loss to a. L = (a - y)², so ∂L/∂a = 2(a - y) = 2(5 - 2) = 6.0
a to h, through the ReLU. ReLU's derivative is 1 if its input was positive and 0 if it was
negative. Our h was 5, positive, so ∂a/∂h = 1.0 and therefore
∂L/∂h = 6.0 × 1.0 = 6.0
h to w and b. Since h = w·x + b:
∂h/∂w = x = 2.0 → ∂L/∂w = 6.0 × 2.0 = 12.0
∂h/∂b = 1.0 → ∂L/∂b = 6.0 × 1.0 = 6.0
So by hand: ∂L/∂w = 12 and ∂L/∂b = 6.
Now ask PyTorch the same question:
x = torch.tensor(2.0)
w = torch.tensor(3.0, requires_grad=True)
b = torch.tensor(-1.0, requires_grad=True)
y = torch.tensor(2.0)
h = w * x + b
a = torch.relu(h)
loss = (a - y) ** 2
loss.backward()
print(f"loss = {loss.item()}") # 9.0
print(f"dL/dw = {w.grad.item()}") # 12.0
print(f"dL/db = {b.grad.item()}") # 6.0
Exact match, all three. Autograd traced the same five operations and applied the same chain rule. It just did it without asking you.
If you read nothing else here, run that block and compare it to the hand calculation above. The whole point of the exercise is the moment the numbers agree.
What happens when the ReLU is closed
Change one number. Keep w = 3 and b = -1, but feed x = -1.0. Now h = 3·(-1) + (-1) = -4, which is negative, so ReLU outputs 0 and its local derivative is also 0.
x = torch.tensor(-1.0)
w = torch.tensor(3.0, requires_grad=True)
b = torch.tensor(-1.0, requires_grad=True)
y = torch.tensor(2.0)
h = w * x + b # -4.0
a = torch.relu(h) # 0.0
loss = (a - y) ** 2 # (0 - 2)² = 4.0
loss.backward()
print(w.grad.item()) # 0.0
print(b.grad.item()) # 0.0
The loss is 4, so the network is wrong, and yet both gradients are exactly zero. No learning
signal reaches w or b at all.
This follows straight from the hand calculation. ∂a/∂h = 0, and every gradient behind that node gets multiplied by zero on its way through. The ReLU is a gate, and this one is shut.
That is the dying-ReLU problem, and it is worth meeting it here rather than three months later in a model with fifty layers. When a neuron's input is negative for every example in your data, its gradient is permanently zero and it never learns again. LeakyReLU exists because it lets a small gradient through instead.
More generally: when a model stops learning and you cannot see why, the question to ask is what is multiplying the gradient by zero on the way back.
Three mechanics that catch nearly everyone
1. Only leaves keep their .grad
Intermediate results do not store their gradient. It is computed, used to keep the chain going, and thrown away, because keeping every intermediate gradient in a real model would be enormous. If you want one, ask before calling backward:
x = torch.tensor(2.0, requires_grad=True)
y = x ** 3
y.retain_grad() # keep y's gradient too
z = y * 5
z.backward()
print(x.grad) # tensor(60.) dz/dx = 15x² = 60
print(y.grad) # tensor(5.) dz/dy = 5
Without retain_grad(), y.grad is None and you get a warning rather than an error, which is why people spend twenty minutes confused by it.
2. Gradients accumulate, they do not overwrite
This one causes more silently broken training loops than anything else in PyTorch.
x = torch.tensor(3.0, requires_grad=True)
for i in range(3):
y = x ** 2
y.backward()
print(f"step {i} grad = {x.grad.item()}")
step 0 grad = 6.0
step 1 grad = 12.0
step 2 grad = 18.0
The gradient of x² at x = 3 is 6, every time. It reads 12 and then 18 because .grad is added to, not replaced.
This is deliberate. It is what lets you accumulate gradients over several mini-batches and take one larger step, which is how people train with an effective batch size their GPU cannot hold.
But it means that unless you clear it, every step of your training loop is stepping on a sum of all previous gradients, and your loss curve will do something strange that is very hard to diagnose from the outside.
The fix is x.grad.zero_(), or in a real loop, optimizer.zero_grad() before every loss.backward(). That line is not boilerplate. It is load-bearing.
3. no_grad and detach — switching the recorder off
During evaluation you do not need gradients, and building the graph costs both time and memory.
x = torch.tensor(3.0, requires_grad=True)
with torch.no_grad():
y = x ** 2
print(y.requires_grad) # False
Inside the block no graph is built. This is why every evaluation loop you have ever copied is wrapped in torch.no_grad().
.detach() is the narrower tool. It gives you a tensor that shares the same data but has no connection to the graph:
x = torch.tensor(3.0, requires_grad=True)
y = x ** 2
y_val = y.detach()
print(y_val) # tensor(9.)
print(y_val.requires_grad) # False
Use it when you want a value without dragging its history along — logging, or using a model's output as a target that should not be differentiated through. It is the mechanism behind stop-gradient tricks and target networks, and it is also, occasionally, the reason your gradient is unexpectedly zero.
One call, every parameter
Everything so far had one or two parameters so the arithmetic stayed visible. Nothing changes at scale.
w = torch.tensor(2.0, requires_grad=True)
b = torch.tensor(1.0, requires_grad=True)
x = torch.tensor(3.0)
y = w * x + b # 2·3 + 1 = 7
loss = (y - 5) ** 2 # (7 - 5)² = 4
loss.backward()
print(w.grad) # tensor(12.)
print(b.grad) # tensor(4.)
One backward() filled in both. In a model with eleven million parameters, the same single call fills in all eleven million, because the graph reaches every one of them and the chain rule multiplies along every path. Nothing about the mechanism is different. There are just more nodes.
The summary worth keeping
-
requires_grad=Truemarks a tensor as something you want the gradient of. Model parameters get it automatically; your data does not need it. - The forward pass computes the answer and records a graph as a side effect. Each node stores what it needs to reverse itself later.
-
.backward()walks that graph in reverse, multiplying the incoming gradient by each node's local derivative. That is the chain rule, and it is all backpropagation is. - Gradients land in
.gradon leaf tensors only. Useretain_grad()for intermediates. - Gradients accumulate.
zero_grad()before every backward, always. -
torch.no_grad()for evaluation,.detach()for a single value without its history. - A zero gradient means something on the path back multiplied by zero. A closed ReLU is the usual suspect.
Try this before you close the tab
Take x = 4.0 with requires_grad=True, compute y = x**3 - 2*x, call backward(), and print
x.grad. Then work out 3x² - 2 at x = 4 on paper and check whether they agree. It takes
about ninety seconds and it is worth more than rereading this article.
Then the harder version: put a second weight after the ReLU in the tiny network above, w2 =, and write out the full chain rule for
0.5∂L/∂w before you run it.
This is one chapter's worth of an idea from my book, PyTorch From Ground Up, which builds
everything from tensors upward so nothing stays vague. If it helped: 8 chapters are free, no email required,
there's a free one-page tensor cheat-sheet here, every example runs
in the companion notebooks on GitHub, and
the full book is on Leanpub or in
paperback and Kindle on Amazon.
More in this series
Shape mechanics, the part that has to be solid before any of this makes sense:
Top comments (0)