DEV Community

Venus-Kennedy
Venus-Kennedy

Posted on

A Beginner’s Guide to Regression and Regularization

Machine learning is often used to make predictions from existing data. For example, a business might want to predict sales, a bank might estimate loan repayment amounts, or a real estate company might predict house prices.

One of the most important techniques for making these types of predictions is regression.

Regression is a supervised machine learning technique used to predict a continuous numerical value based on one or more input variables.

However, building a model that fits the training data extremely well does not always mean that the model will perform well on new data. This problem is closely related to overfitting.

This is where regularization becomes useful.

Regularization adds a penalty to a model's complexity, encouraging it to focus on the most important patterns rather than fitting every small detail or noise in the training data.

In this article, we will explore regression, overfitting, regularization, and the most common regularization techniques used in machine learning.

1. What Is Regression?

Regression is a machine learning approach used to predict a continuous numerical outcome.

Examples of continuous values include:

  • House prices
  • Monthly sales
  • Customer spending
  • Temperature
  • Salary
  • Product demand
  • Stock prices
  • Delivery time

For example, suppose we want to predict the price of a house based on its size.

Our dataset might look like this:

House Size (m²) Price
50 3,500,000
70 4,800,000
90 6,000,000
120 8,200,000
150 10,000,000

A regression model attempts to learn the relationship between house size and price.

Once trained, we could provide the model with a new house size and ask it to estimate the price.

2. Simple Linear Regression

One of the simplest forms of regression is linear regression.

Linear regression attempts to find a straight-line relationship between an input variable and an output variable.

The basic equation is:

ŷ = b₀ + b₁x
Enter fullscreen mode Exit fullscreen mode

Where:

  • ŷ = predicted value
  • b₀ = intercept
  • b₁ = coefficient or slope
  • x = input variable

For example, if we are predicting house prices based on house size, the model might learn that larger houses generally have higher prices.

The model attempts to find the line that best represents the relationship between the input and output values.

3. Multiple Linear Regression

Real-world problems usually involve more than one variable.

For example, house prices may depend on:

  • House size
  • Number of bedrooms
  • Location
  • Age of the property
  • Number of bathrooms
  • Parking spaces

A model can use multiple features:

ŷ = b₀ + b₁x₁ + b₂x₂ + b₃x₃ + ... + bₙxₙ
Enter fullscreen mode Exit fullscreen mode

Here, each x represents a different feature.

For example:

House Price =
    Intercept
    + Size coefficient × Size
    + Bedroom coefficient × Bedrooms
    + Location coefficient × Location score
Enter fullscreen mode Exit fullscreen mode

This is known as multiple linear regression.

4. How Does a Regression Model Learn?

During training, the model compares its predictions with the actual values in the dataset.

The difference between the predicted value and the actual value is called the error or residual.

For example:

Actual price     = 5,000,000
Predicted price  = 4,700,000

Error = 300,000
Enter fullscreen mode Exit fullscreen mode

A regression algorithm tries to find model parameters that minimize the overall prediction error.

One commonly used measure is Mean Squared Error (MSE).

The formula is:

MSE = (1/n) Σ(yᵢ - ŷᵢ)²
Enter fullscreen mode Exit fullscreen mode

Where:

  • n = number of observations
  • yᵢ = actual value
  • ŷᵢ = predicted value

Squaring the errors makes larger errors contribute more heavily to the overall loss.

5. What Is Overfitting?

One of the biggest challenges in machine learning is overfitting.

Overfitting occurs when a model learns the training data too closely, including random noise and patterns that do not generalize to new data.

Imagine a student who memorizes every answer from a practice exam but does not understand the underlying concepts.

The student may perform extremely well on the practice questions but struggle with new questions.

A machine learning model can behave similarly.

Underfitting

The model is too simple and fails to capture important patterns.

Training performance → Poor
Testing performance   → Poor
Enter fullscreen mode Exit fullscreen mode

Good fit

The model captures meaningful patterns and generalizes well.

Training performance → Good
Testing performance   → Good
Enter fullscreen mode Exit fullscreen mode

Overfitting

The model performs extremely well on training data but poorly on unseen data.

Training performance → Very good
Testing performance   → Poor
Enter fullscreen mode Exit fullscreen mode

6. Why Does Overfitting Happen?

Overfitting can happen for several reasons.

Too many features

A dataset may contain many variables, some of which have little useful information.

Too complex a model

A highly flexible model may capture noise instead of meaningful relationships.

Too little training data

With a small dataset, the model may struggle to distinguish real patterns from random variation.

Noisy data

Errors or unusual observations can cause a model to learn patterns that do not generalize.

7. What Is Regularization?

Regularization is a technique used to reduce overfitting by adding a penalty for model complexity.

Instead of asking the model only to minimize prediction error, regularization also encourages the model to keep its coefficients relatively small.

Conceptually:

Total Objective =
Prediction Error
+
Complexity Penalty
Enter fullscreen mode Exit fullscreen mode

The model therefore has to balance two goals:

  1. Fit the training data well.
  2. Avoid becoming unnecessarily complex.

8. Why Penalize Large Coefficients?

Consider a regression model with several features.

Without regularization, the model might assign extremely large coefficients to certain variables to fit the training data more closely.

For example:

ŷ = 5 + 20x₁ + 150x₂ - 90x₃
Enter fullscreen mode Exit fullscreen mode

A regularized model may prefer smaller coefficients:

ŷ = 5 + 8x₁ + 12x₂ - 7x₃
Enter fullscreen mode Exit fullscreen mode

The exact values depend on the data, but the general idea is that regularization discourages unnecessarily large coefficients.

This can make the model less sensitive to noise.

9. The Regularization Parameter

Regularization introduces a parameter commonly represented by λ (lambda).

Lambda controls how strongly the model is penalized for complexity.

Small λ

The penalty is weak.

The model focuses more on fitting the training data.

Small λ → Less regularization
Enter fullscreen mode Exit fullscreen mode

Large λ

The penalty is stronger.

The model is pushed toward simpler coefficients.

Large λ → More regularization
Enter fullscreen mode Exit fullscreen mode

However, making the penalty too strong can cause underfitting.

Therefore, the goal is to find an appropriate level of regularization.

10. Ridge Regression
**
One of the most common regularization techniques is **Ridge Regression
.

Ridge Regression adds a penalty based on the squared values of the model coefficients.

Conceptually, its objective function is:

Loss =
MSE
+
λ Σβⱼ²
Enter fullscreen mode Exit fullscreen mode

The additional term penalizes large coefficients.

Ridge Regression tends to:

  • Reduce the size of coefficients
  • Make models less sensitive to noise
  • Help reduce overfitting
  • Keep all features in the model, although their coefficients may become very small

For example, suppose a model has:

β₁ = 10
β₂ = 8
β₃ = 15
Enter fullscreen mode Exit fullscreen mode

Ridge may shrink them toward smaller values.

Importantly, Ridge generally does not force coefficients exactly to zero.

11. Lasso Regression

Another popular technique is Lasso Regression.

Lasso stands for:

Least Absolute Shrinkage and Selection Operator

Unlike Ridge, Lasso uses the absolute values of coefficients in its penalty.

Conceptually:

Loss =
MSE
+
λ Σ|βⱼ|
Enter fullscreen mode Exit fullscreen mode

Because of this penalty, Lasso can reduce some coefficients all the way to zero.

This makes Lasso particularly useful when feature selection is important.

For example:

Before Lasso:

Feature A → 8.2
Feature B → 0.4
Feature C → 6.7
Feature D → 0.1

After Lasso:

Feature A → 7.8
Feature B → 0
Feature C → 6.3
Feature D → 0
Enter fullscreen mode Exit fullscreen mode

Features B and D effectively become excluded from the model.

*12. Ridge vs. Lasso
*

The key difference is the type of penalty they use.

Feature Ridge Lasso
Penalty Squared coefficients Absolute coefficients
Shrinks coefficients Yes Yes
Can make coefficients exactly zero Generally no Yes
Performs feature selection Limited Yes
Useful for correlated features Often useful Can select among correlated features

A simple way to remember this is:

Ridge shrinks. Lasso can shrink and select.

13. Elastic Net

There is also a third technique called Elastic Net.

Elastic Net combines the penalties used by Ridge and Lasso.

Conceptually:

Loss =
MSE
+
λ₁ Σ|βⱼ|
+
λ₂ Σβⱼ²
Enter fullscreen mode Exit fullscreen mode

Elastic Net therefore combines:

  • Lasso's ability to perform feature selection
  • Ridge's coefficient-shrinking behavior

This can be useful when a dataset contains many correlated features.

14. Why Feature Scaling Matters

Regularization is sensitive to the scale of features.

Imagine a dataset containing:

Age: 18–70
Income: 20,000–500,000
Enter fullscreen mode Exit fullscreen mode

The variables operate on very different scales.

If regularization is applied directly, features with larger numerical scales can have an inappropriate influence on the penalty.

For this reason, features are often standardized before applying regularized regression.

A common standardization approach is:

z = (x - μ) / σ
Enter fullscreen mode Exit fullscreen mode

Where:

  • x = original value
  • μ = mean
  • σ = standard deviation

After standardization, features are placed on comparable scales.

15. Regression Evaluation Metrics

After training a regression model, we need to determine how well it performs.

Several evaluation metrics are commonly used.

Mean Absolute Error (MAE)

MAE measures the average absolute difference between predictions and actual values.

MAE = (1/n) Σ|yᵢ - ŷᵢ|
Enter fullscreen mode Exit fullscreen mode

A lower MAE generally indicates smaller average prediction errors.


Mean Squared Error (MSE)

MSE calculates the average squared prediction error.

MSE = (1/n) Σ(yᵢ - ŷᵢ)²
Enter fullscreen mode Exit fullscreen mode

Because the errors are squared, larger errors receive greater weight.

Root Mean Squared Error (RMSE)

RMSE is the square root of MSE.

RMSE = √MSE
Enter fullscreen mode Exit fullscreen mode

RMSE is useful because it is expressed in the same units as the target variable.

For example, if you are predicting house prices in Kenyan shillings, RMSE is also expressed in Kenyan shillings.

*R² Score
*

R², or the coefficient of determination, measures how much of the variation in the target variable is explained by the model.

Its value is often interpreted in relation to the dataset and modeling context rather than as a universal measure of model quality.

16. Training and Testing Data

To determine whether a regression model generalizes well, we usually divide the dataset into training and testing sets.

For example:

Dataset
   |
   ├── Training Data → Used to train the model
   |
   └── Testing Data → Used to evaluate the model
Enter fullscreen mode Exit fullscreen mode

A model should not be judged only by how well it performs on training data.

A model that performs extremely well on training data but poorly on testing data may be overfitting.

17. Cross-Validation and Regularization

Instead of relying on one train-test split, machine learning practitioners often use cross-validation.

In k-fold cross-validation, the dataset is divided into several sections called folds.

For example, with five-fold cross-validation:

Fold 1
Fold 2
Fold 3
Fold 4
Fold 5
Enter fullscreen mode Exit fullscreen mode

The model is trained and evaluated multiple times, using different folds for validation.

Cross-validation can help estimate how well a model is likely to generalize and can be useful when selecting a suitable regularization parameter.

18. Choosing the Right Regularization Strength

The regularization parameter should generally not be chosen arbitrarily.

A common approach is to test different values using validation or cross-validation.

For example:

λ = 0.001
λ = 0.01
λ = 0.1
λ = 1
λ = 10
Enter fullscreen mode Exit fullscreen mode

The model can be evaluated using cross-validation to determine which setting provides an appropriate balance between fitting the data and controlling complexity.

In practice, machine learning libraries can automate much of this process.

19. A Simple Python Example

Python's scikit-learn library provides implementations of Ridge and Lasso regression.

For example:

from sklearn.linear_model import Ridge

model = Ridge(alpha=1.0)

model.fit(X_train, y_train)

predictions = model.predict(X_test)
Enter fullscreen mode Exit fullscreen mode

Here, alpha controls the strength of the regularization.

A Lasso model can be created similarly:

from sklearn.linear_model import Lasso

model = Lasso(alpha=1.0)

model.fit(X_train, y_train)

predictions = model.predict(X_test)
Enter fullscreen mode Exit fullscreen mode

The appropriate value of alpha depends on the dataset and should generally be selected using a suitable validation strategy.

20. Regression vs. Classification

Regression is sometimes confused with classification, but they solve different types of prediction problems.

Regression

Predicts a continuous numerical value.

Examples:

House price → 8,500,000
Temperature → 27.5°C
Monthly sales → 450,000
Enter fullscreen mode Exit fullscreen mode

Classification

Predicts a category or class.

Examples:

Email → Spam
Transaction → Fraud
Customer → Churn / No Churn
Enter fullscreen mode Exit fullscreen mode

A simple rule is:

Regression predicts numbers; classification predicts categories.

21. Real-World Applications of Regression

Regression is used across many industries.

Finance

Regression can be used to study relationships between financial variables and estimate numerical outcomes.

Real Estate

Models can estimate property prices based on characteristics such as size, location, and number of rooms.

Marketing

Businesses can use regression to study how advertising spending relates to sales.

Healthcare

Regression models can be used to predict numerical outcomes such as treatment costs or length of hospital stay, depending on the application and available data.

*Retail
*

Companies can use regression to forecast demand and analyze factors affecting sales.

Operations

Organizations can model delivery times, resource requirements, and other continuous outcomes.

22. Common Beginner Mistakes

Mistake 1: Thinking a complex model is automatically better

A more complex model can fit training data extremely well but may generalize poorly.

Mistake 2: Ignoring overfitting

Training performance alone does not tell the whole story.

Testing or validation performance is important.

Mistake 3: Forgetting feature scaling

Regularized regression can be sensitive to differences in feature scales.

Mistake 4: Using too much regularization

A very strong penalty can oversimplify the model and lead to underfitting.

Mistake 5: Choosing lambda arbitrarily

The regularization strength should generally be selected using an appropriate validation strategy rather than simply choosing a value at random.

Mistake 6: Confusing regression with classification

Regression predicts continuous numerical outcomes, while classification predicts categories.

23. A Simple Mental Model
**
Think of regression as trying to **draw a useful mathematical relationship through your data
.

Without regularization:

Fit the data as closely as possible
Enter fullscreen mode Exit fullscreen mode

With regularization:

Fit the data
       +
Keep the model reasonably simple
Enter fullscreen mode Exit fullscreen mode

The goal is not to ignore the data.

The goal is to learn patterns that are useful beyond the training dataset.

*24. Key Takeaways
*

Here are the most important ideas to remember:

  1. Regression is used to predict continuous numerical values.

  2. Linear regression models relationships between input variables and a numerical target.

  3. Overfitting occurs when a model learns the training data too closely and performs poorly on unseen data.

  4. Regularization helps reduce overfitting by penalizing model complexity.

  5. Ridge Regression uses an L2 penalty and shrinks coefficients toward zero.

  6. Lasso Regression uses an L1 penalty and can reduce some coefficients to exactly zero.

  7. Elastic Net combines L1 and L2 regularization.

  8. Feature scaling is important when using regularized regression.

  9. Cross-validation can help select an appropriate regularization strength.

  10. A good model should generalize well to new data rather than simply memorizing the training data.

Regression is one of the fundamental techniques in machine learning and provides a way to predict continuous numerical outcomes from data.

However, a model that fits training data extremely well is not necessarily a good model. It may have learned noise or unnecessary complexity, resulting in poor performance on new data.

Regularization helps address this problem by adding a penalty for model complexity.

The three techniques beginners should understand first are:

  • Ridge Regression — shrinks coefficients using an L2 penalty.
  • Lasso Regression — shrinks coefficients and can eliminate some features using an L1 penalty.
  • Elastic Net — combines Ridge and Lasso penalties.

The central idea is simple:

Regression helps us make predictions, while regularization helps us build models that are less likely to overfit.

Understanding these concepts provides a strong foundation for more advanced machine learning topics, including feature selection, model tuning, cross-validation, and predictive modeling.

Top comments (0)