When you start your journey in Machine Learning, the first algorithm you usually learn is Linear Regression. Linear regression draws a straight line using the equation:
This straight line is amazing for predicting continuous numbers, like the price of a house or a student's score. But what happens if you want to predict a category? For example, predicting whether a student is Placed (1) or Not Placed (0) based on their exam scores?
If you try to fit a straight line to binary data (0 and 1), the line keeps going forever. It will eventually predict impossible values like 1.5 or -0.4. Because a straight line fails to categorize discrete choices, Logistic Regression was introduced.
Historical Note: The mathematical foundation of logistic regression was popularized by statistician David Cox in 1958 to analyze binary datasets, adapting the "logistic curve" first formulated by Pierre François Verhulst in 1838 to model population growth.
Instead of drawing a straight line, Logistic Regression squishes the output into an S-shaped curve that strictly stays between 0 and 1. This allows the algorithm to output a clean probability (like an 85% chance of being placed).
Steps of Logistic Regression (The Training Loop)
To see how the computer learns, let's track a tiny dataset with two inputs (CGPA and IQ) to predict a binary output (Placed: 1, Not Placed: 0).
| Student | CGPA ( ) | IQ ( ) | Actual Placed ( ) |
|---|---|---|---|
| Student 1 | 8.0 | 110 | 1 |
| Student 2 | 5.0 | 90 | 0 |
Step 1: Initialize Weights and Biases
If we have
features, we need
weights plus 1 extra intercept term (called the bias or
). For our 2 features, we need
weights. The computer starts by setting them all to zero:
import numpy as np
# n = number of features (CGPA, IQ)
n = 2
weights = np.zeros(n) # Array: [0.0, 0.0]
bias = 0.0 # Scalar: 0.0
Step 2: Make Predictions & Calculate the Loss
The model calculates a raw linear score ( ) for each student using its current weights:
Since all our weights are currently 0, Student 1 gets a score of . We pass this into the Sigmoid function to turn it into a probability ( ):
Now, we check how bad our guess was using Binary Cross-Entropy (BCE) Loss. The formula for a single row is:
For Student 1 (
):
For Student 2 (
):
Average Total Loss:
Step 3: Update Weights Using Gradient Descent
To lower this error, we compute the derivative (gradient) of the loss with respect to each parameter. The math simplifies beautifully to:
Let's look at Student 1 (Predicted , Actual ):
We calculate the weight update by multiplying this error by the student's feature value and our Learning Rate ( ). Since we want to decrease loss, we subtract the gradient:
Step 4: Run the Training Loop
The computer repeats Steps 2 and 3 thousands of times. Over time, the loss drops from 0.693 down toward 0. Let's assume after training, our optimized parameters become:
Let's verify the calculation for Student 1 (CGPA=8.0, IQ=110):
Find
:
Find
(Sigmoid):
Threshold Check: Since
, the model confidently predicts 1 (Placed). The math works!
Important Concepts Explained Simply
- The Sigmoid Function
The Sigmoid function is the mathematical wizard that squishes any real number (from negative infinity to positive infinity) into a neat window between 0 and 1.
Mathematical Formula:
How to read the curve graph:
If
is a large positive number (like +5),
becomes almost 0, so
.
If
is exactly 0,
, so
.
If
is a large negative number (like -5),
, making the denominator massive, so the output drops close to 0.
- The Logit Function (Log of Odds)
You might hear people talk about the "Logit" function. It is simply the exact mathematical inverse of the Sigmoid function. Instead of turning a score into a probability, it turns a probability back into a straight-line score.
- Binary Cross-Entropy (BCE) vs. Categorical Cross-Entropy (CCE)
These are formulas used to measure error.
BCE (Binary): Used when you only have two choices (0 or 1). It only checks the probability of the single correct class.
CCE (Categorical): Used when you have multiple choices (e.g., predicting if an image is a Cat, Dog, or Bird). It measures the error across a distribution of multiple category classes.
- Softmax Function
While Sigmoid handles binary items, Softmax is the multi-class version of Sigmoid. If you are choosing between Cat, Dog, and Bird, Softmax handles all output channels simultaneously and ensures that all their independent probabilities add up to exactly 1.0 (100%).
Complete Working Code From Scratch
Here is how you write this entire mathematical loop in pure Python using numpy:
import numpy as np
class CleanLogisticRegression:
def __init__(self, lr=0.1, num_iterations=1000):
self.lr = lr
self.num_iterations = num_iterations
import numpy as np
self.weights = None
self.bias = None
def _sigmoid(self, z):
return 1 / (1 + np.exp(-z))
def fit(self, X, y):
num_samples, num_features = X.shape
# Step 1: Initialize parameters to zeros
self.weights = np.zeros(num_features)
self.bias = 0.0
# Step 4: The optimization loop
for _ in range(self.num_iterations):
# Step 2: Linear combination and Sigmoid activation
linear_model = np.dot(X, self.weights) + self.bias
y_predicted = self._sigmoid(linear_model)
# Step 3: Compute gradients
dw = (1 / num_samples) * np.dot(X.T, (y_predicted - y))
db = (1 / num_samples) * np.sum(y_predicted - y)
# Step 3: Update parameters (Gradient Descent)
self.weights -= self.lr * dw
self.bias -= self.lr * db
def predict(self, X):
linear_model = np.dot(X, self.weights) + self.bias
y_predicted = self._sigmoid(linear_model)
# Apply the 0.5 threshold logic
return [1 if i >= 0.5 else 0 for i in y_predicted]
# Quick test with our student data
X_train = np.array([[8.0, 110], [5.0, 90]])
y_train = np.array([1, 0])
scratch_model = CleanLogisticRegression(lr=0.1, num_iterations=5000)
scratch_model.fit(X_train, y_train)
print("Scratch Trained Weights:", scratch_model.weights)
print("Scratch Trained Bias:", scratch_model.bias)
How to Use in Scikit-Learn
In industrial settings, you don't need to write code from scratch. You can use Python's built-in scikit-learn framework:
from sklearn.linear_model import LogisticRegression
# 1. Initialize the model with chosen hyperparameters
# 'penalty' controls regularization to avoid overfitting
# 'C' controls parameter strength (smaller C means stronger regularization)
model = LogisticRegression(penalty='l2', C=1.0, max_iter=1000)
# 2. Train the model using the fit function
model.fit(X_train, y_train)
# 3. View the learned coefficients
print("Sklearn Weights:", model.coef_)
print("Sklearn Intercept:", model.intercept_)
# 4. Predict on new incoming data (CGPA = 7.2, IQ = 105)
new_student = [[7.2, 105]]
prediction = model.predict(new_student)
probability = model.predict_proba(new_student)
print(f"Prediction Class: {prediction}")
print(f"Probabilities (Not Placed vs Placed): {probability}")
Core Hyperparameters Explained:
penalty: Can be set to 'l1' or 'l2'. It adds a mathematical penalty score to the loss function to prevent individual weights from becoming too large, which keeps the model stable and prevents overfitting.
C: Inverse of regularization strength. If you set C to a very small number (like 0.01), you tell the model: "Keep the weights very close to zero." If you make it large (like 100), you tell the model: "Focus entirely on fitting the training points perfectly."
max_iter: This is the maximum number of loops you allow the Gradient Descent algorithm to execute to find your ideal weights.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.