DEV Community

Abhishek Sharma
Abhishek Sharma

Posted on

Logistic Regression

When you start your journey in Machine Learning, the first algorithm you usually learn is Linear Regression. Linear regression draws a straight line using the equation:

y=mx+cy = mx + c

This straight line is amazing for predicting continuous numbers, like the price of a house or a student's score. But what happens if you want to predict a category? For example, predicting whether a student is Placed (1) or Not Placed (0) based on their exam scores?

If you try to fit a straight line to binary data (0 and 1), the line keeps going forever. It will eventually predict impossible values like 1.5 or -0.4. Because a straight line fails to categorize discrete choices, Logistic Regression was introduced.

Historical Note: The mathematical foundation of logistic regression was popularized by statistician David Cox in 1958 to analyze binary datasets, adapting the "logistic curve" first formulated by Pierre François Verhulst in 1838 to model population growth.

Instead of drawing a straight line, Logistic Regression squishes the output into an S-shaped curve that strictly stays between 0 and 1. This allows the algorithm to output a clean probability (like an 85% chance of being placed).


Steps of Logistic Regression (The Training Loop)

To see how the computer learns, let's track a tiny dataset with two inputs (CGPA and IQ) to predict a binary output (Placed: 1, Not Placed: 0).

Student CGPA ( X1X_{1} ) IQ ( X2X_{2} ) Actual Placed ( YY )
Student 1 8.0 110 1
Student 2 5.0 90 0

Step 1: Initialize Weights and Biases

If we have nn features, we need nn weights plus 1 extra intercept term (called the bias or β0\beta_{0} ). For our 2 features, we need 2+1=32 + 1 = 3 weights. The computer starts by setting them all to zero:

import numpy as np

# n = number of features (CGPA, IQ)
n = 2
weights = np.zeros(n)  # Array: [0.0, 0.0]
bias = 0.0             # Scalar: 0.0
Enter fullscreen mode Exit fullscreen mode

Step 2: Make Predictions & Calculate the Loss

The model calculates a raw linear score ( zz ) for each student using its current weights:

z=bias+(w1×CGPA)+(w2×IQ) z = \text{bias} + (w_{1} \times \text{CGPA}) + (w_{2} \times \text{IQ})

Since all our weights are currently 0, Student 1 gets a score of z=0z = 0 . We pass this zz into the Sigmoid function to turn it into a probability ( pp ):

p=11+e−0=0.5(50%) p = \frac{1}{1 + e^{-0}} = 0.5 \quad (50\%)

Now, we check how bad our guess was using Binary Cross-Entropy (BCE) Loss. The formula for a single row is:

Loss=−[Ylog⁡(p)+(1−Y)log⁡(1−p)] \text{Loss} = -[Y \log(p) + (1 - Y) \log(1 - p)]
For Student 1 (

  Y=1,p=0.5Y=1, p=0.5

): 

  Loss=−[1⋅log⁡(0.5)+0]=−(−0.693)=0.693\text{Loss} = -[1 \cdot \log(0.5) + 0] = -(-0.693) = 0.693


For Student 2 (

  Y=0,p=0.5Y=0, p=0.5

): 

  Loss=−[0+1⋅log⁡(1−0.5)]=−(−0.693)=0.693\text{Loss} = -[0 + 1 \cdot \log(1 - 0.5)] = -(-0.693) = 0.693


Average Total Loss: 

  0.693+0.6932=0.693\frac{0.693 + 0.693}{2} = 0.693


Enter fullscreen mode Exit fullscreen mode

Step 3: Update Weights Using Gradient Descent

To lower this error, we compute the derivative (gradient) of the loss with respect to each parameter. The math simplifies beautifully to:

Gradient=Predicted Probability−Actual Outcome \text{Gradient} = \text{Predicted Probability} - \text{Actual Outcome}

Let's look at Student 1 (Predicted p=0.5p = 0.5 , Actual Y=1Y = 1 ):

Error=0.5−1=−0.5\text{Error} = 0.5 - 1 = -0.5

We calculate the weight update by multiplying this error by the student's feature value and our Learning Rate ( lr=0.1lr = 0.1 ). Since we want to decrease loss, we subtract the gradient:

New Parameter=Old Parameter−(lr×Gradient) \text{New Parameter} = \text{Old Parameter} - (lr \times \text{Gradient})

Step 4: Run the Training Loop

The computer repeats Steps 2 and 3 thousands of times. Over time, the loss drops from 0.693 down toward 0. Let's assume after training, our optimized parameters become:



  bias=−15\text{bias} = -15




  w1 (CGPA)=1.5w_1 \text{ (CGPA)} = 1.5




  w2 (IQ)=0.05w_2 \text{ (IQ)} = 0.05


Enter fullscreen mode Exit fullscreen mode

Let's verify the calculation for Student 1 (CGPA=8.0, IQ=110):

Find 

  zz

: 

  z=−15+(1.5×8.0)+(0.05×110)=−15+12+5.5=+2.5z = -15 + (1.5 \times 8.0) + (0.05 \times 110) = -15 + 12 + 5.5 = +2.5


Find 

  pp

 (Sigmoid): 

  p=11+e−2.5=11+0.082=11.082≈0.924(92.4p = \frac{1}{1 + e^{-2.5}} = \frac{1}{1 + 0.082} = \frac{1}{1.082} \approx 0.924 \quad (92.4%)


Threshold Check: Since 

  0.924≥0.50.924 \ge 0.5

, the model confidently predicts 1 (Placed). The math works!
Enter fullscreen mode Exit fullscreen mode

Important Concepts Explained Simply

  1. The Sigmoid Function

The Sigmoid function is the mathematical wizard that squishes any real number (from negative infinity to positive infinity) into a neat window between 0 and 1.

Mathematical Formula:

σ(z)=11+e−z \sigma(z) = \frac{1}{1 + e^{-z}}

How to read the curve graph:

If 

  zz

 is a large positive number (like +5), 

  e−5e^{-5}

 becomes almost 0, so 

  11+0=1\frac{1}{1+0} = 1

.
If 

  zz

 is exactly 0, 

  e0=1e^{0} = 1

, so 

  11+1=0.5\frac{1}{1+1} = 0.5

.
If 

  zz

 is a large negative number (like -5), 

  e−(−5)=e5e^{-(-5)} = e^5

, making the denominator massive, so the output drops close to 0.
Enter fullscreen mode Exit fullscreen mode
  1. The Logit Function (Log of Odds)

You might hear people talk about the "Logit" function. It is simply the exact mathematical inverse of the Sigmoid function. Instead of turning a score into a probability, it turns a probability back into a straight-line score.

Odds=p1−p \text{Odds} = \frac{p}{1 - p}
Logit(p)=log⁡(p1−p)=β0+β1X1+… \text{Logit}(p) = \log \left(\frac{p}{1 - p}\right) = \beta_{0} + \beta_{1}X_{1} + \dots
  1. Binary Cross-Entropy (BCE) vs. Categorical Cross-Entropy (CCE)

These are formulas used to measure error.

BCE (Binary): Used when you only have two choices (0 or 1). It only checks the probability of the single correct class.
CCE (Categorical): Used when you have multiple choices (e.g., predicting if an image is a Cat, Dog, or Bird). It measures the error across a distribution of multiple category classes.
Enter fullscreen mode Exit fullscreen mode
  1. Softmax Function

While Sigmoid handles binary items, Softmax is the multi-class version of Sigmoid. If you are choosing between Cat, Dog, and Bird, Softmax handles all output channels simultaneously and ensures that all their independent probabilities add up to exactly 1.0 (100%).

Softmax(zi)=ezi∑ezj \text{Softmax}(z_{i}) = \frac{e^{z_{i}}}{\sum e^{z_{j}}}



Complete Working Code From Scratch

Here is how you write this entire mathematical loop in pure Python using numpy:

import numpy as np

class CleanLogisticRegression:
    def __init__(self, lr=0.1, num_iterations=1000):
        self.lr = lr
        self.num_iterations = num_iterations
import numpy as np
        self.weights = None
        self.bias = None

    def _sigmoid(self, z):
        return 1 / (1 + np.exp(-z))

    def fit(self, X, y):
        num_samples, num_features = X.shape
        # Step 1: Initialize parameters to zeros
        self.weights = np.zeros(num_features)
        self.bias = 0.0

        # Step 4: The optimization loop
        for _ in range(self.num_iterations):
            # Step 2: Linear combination and Sigmoid activation
            linear_model = np.dot(X, self.weights) + self.bias
            y_predicted = self._sigmoid(linear_model)

            # Step 3: Compute gradients
            dw = (1 / num_samples) * np.dot(X.T, (y_predicted - y))
            db = (1 / num_samples) * np.sum(y_predicted - y)

            # Step 3: Update parameters (Gradient Descent)
            self.weights -= self.lr * dw
            self.bias -= self.lr * db

    def predict(self, X):
        linear_model = np.dot(X, self.weights) + self.bias
        y_predicted = self._sigmoid(linear_model)
        # Apply the 0.5 threshold logic
        return [1 if i >= 0.5 else 0 for i in y_predicted]

# Quick test with our student data
X_train = np.array([[8.0, 110], [5.0, 90]])
y_train = np.array([1, 0])

scratch_model = CleanLogisticRegression(lr=0.1, num_iterations=5000)
scratch_model.fit(X_train, y_train)
print("Scratch Trained Weights:", scratch_model.weights)
print("Scratch Trained Bias:", scratch_model.bias)
Enter fullscreen mode Exit fullscreen mode

How to Use in Scikit-Learn

In industrial settings, you don't need to write code from scratch. You can use Python's built-in scikit-learn framework:

from sklearn.linear_model import LogisticRegression

# 1. Initialize the model with chosen hyperparameters
# 'penalty' controls regularization to avoid overfitting
# 'C' controls parameter strength (smaller C means stronger regularization)
model = LogisticRegression(penalty='l2', C=1.0, max_iter=1000)

# 2. Train the model using the fit function
model.fit(X_train, y_train)

# 3. View the learned coefficients
print("Sklearn Weights:", model.coef_)
print("Sklearn Intercept:", model.intercept_)

# 4. Predict on new incoming data (CGPA = 7.2, IQ = 105)
new_student = [[7.2, 105]]
prediction = model.predict(new_student)
probability = model.predict_proba(new_student)

print(f"Prediction Class: {prediction}")
print(f"Probabilities (Not Placed vs Placed): {probability}")
Enter fullscreen mode Exit fullscreen mode

Core Hyperparameters Explained:

penalty: Can be set to 'l1' or 'l2'. It adds a mathematical penalty score to the loss function to prevent individual weights from becoming too large, which keeps the model stable and prevents overfitting.
C: Inverse of regularization strength. If you set C to a very small number (like 0.01), you tell the model: "Keep the weights very close to zero." If you make it large (like 100), you tell the model: "Focus entirely on fitting the training points perfectly."
max_iter: This is the maximum number of loops you allow the Gradient Descent algorithm to execute to find your ideal weights. 
Enter fullscreen mode Exit fullscreen mode

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.