DEV Community

Cover image for I Built a Tiny Neural Network From Scratch in Python — No PyTorch
Utsav D
Utsav D

Posted on

I Built a Tiny Neural Network From Scratch in Python — No PyTorch

I Built a Tiny Neural Network From Scratch in Python — No PyTorch

I didn't want to just call model.fit() and say I understood neural networks.

So I decided to build one myself.

No PyTorch.
No TensorFlow.
No Keras.

Just Python + NumPy + the mathematics behind a neural network.

The goal wasn't to build a production-ready deep learning framework. The goal was to understand what actually happens inside a neural network during training.

By the end, we will have a small neural network that can learn a binary classification problem using:

  • Forward propagation
  • ReLU activation
  • Sigmoid activation
  • Binary cross-entropy loss
  • Backpropagation
  • Gradient descent
  • NumPy matrix operations

1. What Are We Actually Building?

Our network will have three parts:

Input Layer
    ↓
Hidden Layer
    ↓
Output Layer
Enter fullscreen mode Exit fullscreen mode

For this example:

3 input features
      ↓
4 hidden neurons
      ↓
1 output neuron
Enter fullscreen mode Exit fullscreen mode

The output represents the probability of belonging to class 1.

For example:

0.02 → Class 0
0.91 → Class 1
Enter fullscreen mode Exit fullscreen mode

2. The Mathematics Behind a Neuron

A neuron first calculates a weighted sum:

$$
z = XW + b
$$

where:

  • X = input
  • W = weights
  • b = bias
  • z = pre-activation value

Then an activation function is applied.

For the hidden layer we'll use ReLU:

$$
ReLU(x) = \max(0,x)
$$

For the output layer we'll use sigmoid:

$$
\sigma(x) = \frac{1}{1+e^{-x}}
$$

The sigmoid converts the output into a value between 0 and 1.


3. Creating Some Data

We'll create a small synthetic binary classification dataset.

import numpy as np

np.random.seed(42)

n_samples = 200

X = np.random.randn(n_samples, 3)

y = (
    (X[:, 0] + X[:, 1] - X[:, 2]) > 0
).astype(int).reshape(-1, 1)

print(X.shape)
print(y.shape)
Enter fullscreen mode Exit fullscreen mode

The shapes are:

X → (200, 3)
y → (200, 1)
Enter fullscreen mode Exit fullscreen mode

So every sample contains three features.


4. Activation Functions

Let's implement ReLU and sigmoid ourselves.

def relu(x):
    return np.maximum(0, x)


def relu_derivative(x):
    return (x > 0).astype(float)


def sigmoid(x):
    return 1 / (1 + np.exp(-x))
Enter fullscreen mode Exit fullscreen mode

The derivative of ReLU is important during backpropagation.

It is:

1 when x > 0
0 when x <= 0
Enter fullscreen mode Exit fullscreen mode

5. Building the Neural Network

Now we can create the actual network.

class NeuralNetwork:

    def __init__(self, input_size, hidden_size, output_size):

        self.W1 = np.random.randn(input_size, hidden_size) * 0.1
        self.b1 = np.zeros((1, hidden_size))

        self.W2 = np.random.randn(hidden_size, output_size) * 0.1
        self.b2 = np.zeros((1, output_size))
Enter fullscreen mode Exit fullscreen mode

The architecture is:

3 → 4 → 1
Enter fullscreen mode Exit fullscreen mode

Therefore:

W1 = 3 × 4
b1 = 1 × 4

W2 = 4 × 1
b2 = 1 × 1
Enter fullscreen mode Exit fullscreen mode

6. Forward Propagation

Now we need to pass the input through the network.

def forward(self, X):

    self.z1 = X @ self.W1 + self.b1
    self.a1 = relu(self.z1)

    self.z2 = self.a1 @ self.W2 + self.b2
    self.a2 = sigmoid(self.z2)

    return self.a2
Enter fullscreen mode Exit fullscreen mode

Mathematically:

$$
Z_1 = XW_1+b_1
$$

$$
A_1 = ReLU(Z_1)
$$

$$
Z_2 = A_1W_2+b_2
$$

$$
A_2 = Sigmoid(Z_2)
$$

And A2 becomes our prediction.


7. Measuring the Error

A model needs to know how wrong its predictions are.

For binary classification, we'll use binary cross-entropy:

$$
L = -\frac{1}{m}\sum
[y\log(\hat y)+(1-y)\log(1-\hat y)]
$$

In Python:

def binary_cross_entropy(y, y_pred):

    epsilon = 1e-8

    y_pred = np.clip(
        y_pred,
        epsilon,
        1 - epsilon
    )

    return -np.mean(
        y * np.log(y_pred)
        + (1 - y) * np.log(1 - y_pred)
    )
Enter fullscreen mode Exit fullscreen mode

The clip() prevents numerical problems caused by taking the logarithm of zero.


8. The Part That Usually Gets Hidden: Backpropagation

This is where things get interesting.

During backpropagation, we calculate how much each parameter contributed to the error.

For the output layer:

$$
dZ_2 = A_2-y
$$

Then:

$$
dW_2 = \frac{A_1^T dZ_2}{m}
$$

$$
db_2 = \frac{\sum dZ_2}{m}
$$

Then we propagate the gradient back into the hidden layer:

$$
dA_1 = dZ_2W_2^T
$$

$$
dZ_1 = dA_1 \odot ReLU'(Z_1)
$$

Finally:

$$
dW_1 = \frac{X^TdZ_1}{m}
$$

and:

$$
db_1 = \frac{\sum dZ_1}{m}
$$

The implementation:

def backward(self, X, y, learning_rate=0.01):

    m = X.shape[0]

    # Output layer gradient
    dz2 = self.a2 - y

    dw2 = (self.a1.T @ dz2) / m
    db2 = np.sum(dz2, axis=0, keepdims=True) / m

    # Hidden layer gradient
    da1 = dz2 @ self.W2.T

    dz1 = da1 * relu_derivative(self.z1)

    dw1 = (X.T @ dz1) / m
    db1 = np.sum(dz1, axis=0, keepdims=True) / m

    # Update parameters
    self.W2 -= learning_rate * dw2
    self.b2 -= learning_rate * db2

    self.W1 -= learning_rate * dw1
    self.b1 -= learning_rate * db1
Enter fullscreen mode Exit fullscreen mode

That's essentially the core of training.


9. Training the Network

Now let's put everything together.

model = NeuralNetwork(
    input_size=3,
    hidden_size=4,
    output_size=1
)

epochs = 2000
learning_rate = 0.05

loss_history = []

for epoch in range(epochs):

    # Forward pass
    predictions = model.forward(X)

    # Calculate loss
    loss = binary_cross_entropy(
        y,
        predictions
    )

    # Backpropagation
    model.backward(
        X,
        y,
        learning_rate
    )

    loss_history.append(loss)

    if epoch % 200 == 0:
        print(
            f"Epoch {epoch}, Loss: {loss:.4f}"
        )
Enter fullscreen mode Exit fullscreen mode

At the beginning, the network has essentially no idea what it is doing.

Its weights are randomly initialized.

But every iteration does this:

Input
  ↓
Prediction
  ↓
Calculate Error
  ↓
Calculate Gradients
  ↓
Update Weights
  ↓
Repeat
Enter fullscreen mode Exit fullscreen mode

Over time, the loss should decrease.


10. Making Predictions

After training, we can use the network to predict new samples.

predictions = model.forward(X)

predicted_classes = (
    predictions >= 0.5
).astype(int)
Enter fullscreen mode Exit fullscreen mode

Let's calculate accuracy:

accuracy = np.mean(
    predicted_classes == y
)

print(
    f"Accuracy: {accuracy * 100:.2f}%"
)
Enter fullscreen mode Exit fullscreen mode

The important part isn't the exact number.

The important part is that we built the learning process ourselves.


11. Looking at the Training Process

We can visualize the loss:

import matplotlib.pyplot as plt

plt.plot(loss_history)

plt.xlabel("Epoch")
plt.ylabel("Loss")
plt.title("Training Loss")

plt.show()
Enter fullscreen mode Exit fullscreen mode

You should see the loss generally decrease as training progresses.

That curve is one of the simplest ways to see learning happening.


12. What Is Actually Happening During One Training Step?

It can be reduced to seven steps:

Step 1 — Input

X
Enter fullscreen mode Exit fullscreen mode

Step 2 — Forward propagation

X → Hidden Layer → Output
Enter fullscreen mode Exit fullscreen mode

Step 3 — Prediction

ŷ
Enter fullscreen mode Exit fullscreen mode

Step 4 — Loss

ŷ compared with y
Enter fullscreen mode Exit fullscreen mode

Step 5 — Backpropagation

Calculate gradients.

Step 6 — Gradient descent

Update the weights:

$$
W := W-\eta\frac{\partial L}{\partial W}
$$

where η is the learning rate.

Step 7 — Repeat

Thousands of times.

That's training.


13. So What Does PyTorch Actually Do?

After building this manually, frameworks like PyTorch become much easier to understand.

When you write:

loss.backward()
Enter fullscreen mode Exit fullscreen mode

PyTorch is performing the gradient calculations for you.

When you write:

optimizer.step()
Enter fullscreen mode Exit fullscreen mode

the parameters are updated.

And when you define:

nn.Linear(3, 4)
Enter fullscreen mode Exit fullscreen mode

you're creating something conceptually similar to:

W = np.random.randn(3, 4)
b = np.zeros((1, 4))
Enter fullscreen mode Exit fullscreen mode

Of course, real frameworks handle much more:

  • Automatic differentiation
  • GPU acceleration
  • Optimizers
  • Memory management
  • Tensor operations
  • Neural-network modules
  • Mixed precision
  • Distributed training

But the fundamental ideas are still the same.


14. Why Build One From Scratch?

You might be wondering:

Why spend time implementing something that PyTorch already does?

Because using a framework and understanding the underlying process are different things.

When you only use:

model.fit(...)
Enter fullscreen mode Exit fullscreen mode

it's easy to think of training as a black box.

Building the network manually forces you to understand:

  • Why weights exist
  • Why biases exist
  • What activation functions do
  • What a gradient represents
  • Why the loss changes
  • How errors move backward
  • Why learning rate matters
  • Why matrix dimensions matter

Once those concepts click, high-level deep-learning code becomes much less mysterious.


15. What I Would Add Next

This implementation is intentionally tiny.

A more serious version could add:

  • Mini-batch gradient descent
  • Multiple hidden layers
  • Softmax for multiclass classification
  • Adam optimizer
  • Dropout
  • Batch normalization
  • L2 regularization
  • Model saving/loading
  • Train/validation/test splits
  • Hyperparameter tuning

The next interesting experiment would be to implement the same architecture twice:

NumPy implementation
        vs
PyTorch implementation
Enter fullscreen mode Exit fullscreen mode

Then compare their training behavior and code complexity.


16. Project Structure

A simple project could look like this:

tiny-neural-network/
│
├── neural_network.py
├── train.py
├── data.py
├── visualize.py
├── requirements.txt
└── README.md
Enter fullscreen mode Exit fullscreen mode

requirements.txt could be as simple as:

numpy
matplotlib
Enter fullscreen mode Exit fullscreen mode

That's it.

No deep-learning framework is required.


17. The Biggest Thing I Learned

The most useful part of this experiment wasn't the final classifier.

It was realizing how much abstraction modern frameworks provide.

A few lines of PyTorch can represent operations that involve:

Matrix multiplication
        ↓
Activation
        ↓
Loss
        ↓
Derivatives
        ↓
Gradient calculation
        ↓
Parameter updates
Enter fullscreen mode Exit fullscreen mode

When you understand those operations individually, frameworks stop feeling like magic.

They become tools.


Final Takeaway

You don't need to build every machine-learning model from scratch.

In fact, for real projects, you usually shouldn't.

PyTorch and TensorFlow exist for very good reasons.

But building a tiny neural network once is an excellent way to understand what happens underneath the APIs.

The next time you see:

loss.backward()
optimizer.step()
Enter fullscreen mode Exit fullscreen mode

you'll have a much better idea of what those two lines actually represent.

The framework is the abstraction.
The mathematics is what makes it work.

Top comments (0)