I Built a Tiny Neural Network From Scratch in Python — No PyTorch
I didn't want to just call model.fit() and say I understood neural networks.
So I decided to build one myself.
No PyTorch.
No TensorFlow.
No Keras.
Just Python + NumPy + the mathematics behind a neural network.
The goal wasn't to build a production-ready deep learning framework. The goal was to understand what actually happens inside a neural network during training.
By the end, we will have a small neural network that can learn a binary classification problem using:
- Forward propagation
- ReLU activation
- Sigmoid activation
- Binary cross-entropy loss
- Backpropagation
- Gradient descent
- NumPy matrix operations
1. What Are We Actually Building?
Our network will have three parts:
Input Layer
↓
Hidden Layer
↓
Output Layer
For this example:
3 input features
↓
4 hidden neurons
↓
1 output neuron
The output represents the probability of belonging to class 1.
For example:
0.02 → Class 0
0.91 → Class 1
2. The Mathematics Behind a Neuron
A neuron first calculates a weighted sum:
$$
z = XW + b
$$
where:
-
X= input -
W= weights -
b= bias -
z= pre-activation value
Then an activation function is applied.
For the hidden layer we'll use ReLU:
$$
ReLU(x) = \max(0,x)
$$
For the output layer we'll use sigmoid:
$$
\sigma(x) = \frac{1}{1+e^{-x}}
$$
The sigmoid converts the output into a value between 0 and 1.
3. Creating Some Data
We'll create a small synthetic binary classification dataset.
import numpy as np
np.random.seed(42)
n_samples = 200
X = np.random.randn(n_samples, 3)
y = (
(X[:, 0] + X[:, 1] - X[:, 2]) > 0
).astype(int).reshape(-1, 1)
print(X.shape)
print(y.shape)
The shapes are:
X → (200, 3)
y → (200, 1)
So every sample contains three features.
4. Activation Functions
Let's implement ReLU and sigmoid ourselves.
def relu(x):
return np.maximum(0, x)
def relu_derivative(x):
return (x > 0).astype(float)
def sigmoid(x):
return 1 / (1 + np.exp(-x))
The derivative of ReLU is important during backpropagation.
It is:
1 when x > 0
0 when x <= 0
5. Building the Neural Network
Now we can create the actual network.
class NeuralNetwork:
def __init__(self, input_size, hidden_size, output_size):
self.W1 = np.random.randn(input_size, hidden_size) * 0.1
self.b1 = np.zeros((1, hidden_size))
self.W2 = np.random.randn(hidden_size, output_size) * 0.1
self.b2 = np.zeros((1, output_size))
The architecture is:
3 → 4 → 1
Therefore:
W1 = 3 × 4
b1 = 1 × 4
W2 = 4 × 1
b2 = 1 × 1
6. Forward Propagation
Now we need to pass the input through the network.
def forward(self, X):
self.z1 = X @ self.W1 + self.b1
self.a1 = relu(self.z1)
self.z2 = self.a1 @ self.W2 + self.b2
self.a2 = sigmoid(self.z2)
return self.a2
Mathematically:
$$
Z_1 = XW_1+b_1
$$
$$
A_1 = ReLU(Z_1)
$$
$$
Z_2 = A_1W_2+b_2
$$
$$
A_2 = Sigmoid(Z_2)
$$
And A2 becomes our prediction.
7. Measuring the Error
A model needs to know how wrong its predictions are.
For binary classification, we'll use binary cross-entropy:
$$
L = -\frac{1}{m}\sum
[y\log(\hat y)+(1-y)\log(1-\hat y)]
$$
In Python:
def binary_cross_entropy(y, y_pred):
epsilon = 1e-8
y_pred = np.clip(
y_pred,
epsilon,
1 - epsilon
)
return -np.mean(
y * np.log(y_pred)
+ (1 - y) * np.log(1 - y_pred)
)
The clip() prevents numerical problems caused by taking the logarithm of zero.
8. The Part That Usually Gets Hidden: Backpropagation
This is where things get interesting.
During backpropagation, we calculate how much each parameter contributed to the error.
For the output layer:
$$
dZ_2 = A_2-y
$$
Then:
$$
dW_2 = \frac{A_1^T dZ_2}{m}
$$
$$
db_2 = \frac{\sum dZ_2}{m}
$$
Then we propagate the gradient back into the hidden layer:
$$
dA_1 = dZ_2W_2^T
$$
$$
dZ_1 = dA_1 \odot ReLU'(Z_1)
$$
Finally:
$$
dW_1 = \frac{X^TdZ_1}{m}
$$
and:
$$
db_1 = \frac{\sum dZ_1}{m}
$$
The implementation:
def backward(self, X, y, learning_rate=0.01):
m = X.shape[0]
# Output layer gradient
dz2 = self.a2 - y
dw2 = (self.a1.T @ dz2) / m
db2 = np.sum(dz2, axis=0, keepdims=True) / m
# Hidden layer gradient
da1 = dz2 @ self.W2.T
dz1 = da1 * relu_derivative(self.z1)
dw1 = (X.T @ dz1) / m
db1 = np.sum(dz1, axis=0, keepdims=True) / m
# Update parameters
self.W2 -= learning_rate * dw2
self.b2 -= learning_rate * db2
self.W1 -= learning_rate * dw1
self.b1 -= learning_rate * db1
That's essentially the core of training.
9. Training the Network
Now let's put everything together.
model = NeuralNetwork(
input_size=3,
hidden_size=4,
output_size=1
)
epochs = 2000
learning_rate = 0.05
loss_history = []
for epoch in range(epochs):
# Forward pass
predictions = model.forward(X)
# Calculate loss
loss = binary_cross_entropy(
y,
predictions
)
# Backpropagation
model.backward(
X,
y,
learning_rate
)
loss_history.append(loss)
if epoch % 200 == 0:
print(
f"Epoch {epoch}, Loss: {loss:.4f}"
)
At the beginning, the network has essentially no idea what it is doing.
Its weights are randomly initialized.
But every iteration does this:
Input
↓
Prediction
↓
Calculate Error
↓
Calculate Gradients
↓
Update Weights
↓
Repeat
Over time, the loss should decrease.
10. Making Predictions
After training, we can use the network to predict new samples.
predictions = model.forward(X)
predicted_classes = (
predictions >= 0.5
).astype(int)
Let's calculate accuracy:
accuracy = np.mean(
predicted_classes == y
)
print(
f"Accuracy: {accuracy * 100:.2f}%"
)
The important part isn't the exact number.
The important part is that we built the learning process ourselves.
11. Looking at the Training Process
We can visualize the loss:
import matplotlib.pyplot as plt
plt.plot(loss_history)
plt.xlabel("Epoch")
plt.ylabel("Loss")
plt.title("Training Loss")
plt.show()
You should see the loss generally decrease as training progresses.
That curve is one of the simplest ways to see learning happening.
12. What Is Actually Happening During One Training Step?
It can be reduced to seven steps:
Step 1 — Input
X
Step 2 — Forward propagation
X → Hidden Layer → Output
Step 3 — Prediction
ŷ
Step 4 — Loss
ŷ compared with y
Step 5 — Backpropagation
Calculate gradients.
Step 6 — Gradient descent
Update the weights:
$$
W := W-\eta\frac{\partial L}{\partial W}
$$
where η is the learning rate.
Step 7 — Repeat
Thousands of times.
That's training.
13. So What Does PyTorch Actually Do?
After building this manually, frameworks like PyTorch become much easier to understand.
When you write:
loss.backward()
PyTorch is performing the gradient calculations for you.
When you write:
optimizer.step()
the parameters are updated.
And when you define:
nn.Linear(3, 4)
you're creating something conceptually similar to:
W = np.random.randn(3, 4)
b = np.zeros((1, 4))
Of course, real frameworks handle much more:
- Automatic differentiation
- GPU acceleration
- Optimizers
- Memory management
- Tensor operations
- Neural-network modules
- Mixed precision
- Distributed training
But the fundamental ideas are still the same.
14. Why Build One From Scratch?
You might be wondering:
Why spend time implementing something that PyTorch already does?
Because using a framework and understanding the underlying process are different things.
When you only use:
model.fit(...)
it's easy to think of training as a black box.
Building the network manually forces you to understand:
- Why weights exist
- Why biases exist
- What activation functions do
- What a gradient represents
- Why the loss changes
- How errors move backward
- Why learning rate matters
- Why matrix dimensions matter
Once those concepts click, high-level deep-learning code becomes much less mysterious.
15. What I Would Add Next
This implementation is intentionally tiny.
A more serious version could add:
- Mini-batch gradient descent
- Multiple hidden layers
- Softmax for multiclass classification
- Adam optimizer
- Dropout
- Batch normalization
- L2 regularization
- Model saving/loading
- Train/validation/test splits
- Hyperparameter tuning
The next interesting experiment would be to implement the same architecture twice:
NumPy implementation
vs
PyTorch implementation
Then compare their training behavior and code complexity.
16. Project Structure
A simple project could look like this:
tiny-neural-network/
│
├── neural_network.py
├── train.py
├── data.py
├── visualize.py
├── requirements.txt
└── README.md
requirements.txt could be as simple as:
numpy
matplotlib
That's it.
No deep-learning framework is required.
17. The Biggest Thing I Learned
The most useful part of this experiment wasn't the final classifier.
It was realizing how much abstraction modern frameworks provide.
A few lines of PyTorch can represent operations that involve:
Matrix multiplication
↓
Activation
↓
Loss
↓
Derivatives
↓
Gradient calculation
↓
Parameter updates
When you understand those operations individually, frameworks stop feeling like magic.
They become tools.
Final Takeaway
You don't need to build every machine-learning model from scratch.
In fact, for real projects, you usually shouldn't.
PyTorch and TensorFlow exist for very good reasons.
But building a tiny neural network once is an excellent way to understand what happens underneath the APIs.
The next time you see:
loss.backward()
optimizer.step()
you'll have a much better idea of what those two lines actually represent.
The framework is the abstraction.
The mathematics is what makes it work.
Top comments (0)