DEV Community

Amailuk Joseph
Amailuk Joseph

Posted on

The Detective's Guide to Regression & Regularization

The Detective's Guide to Regression & Regularization

Phase 1: The Lemonade Stand Mystery


πŸ•΅οΈ Detective's Notebook

(This is a brand new case β€” nothing to recap yet. But by the end of this phase, you'll have solved your first real mystery in data science, even if nobody's told you that's what you're doing.)


The Problem

Meet Zara. She's 13, and every summer she runs a lemonade stand outside her house.

Some days she sells 5 cups. Some days she sells 40 cups. And every single morning, she faces the same annoying challenge:

How many cups should she make?

Make too few, and she runs out by 11 a.m., watching thirsty customers walk away. Make too many, and she's pouring warm, wasted lemonade down the drain by evening.

Zara doesn't have a crystal ball. But she does have something almost as good: a notebook. For the last 14 days, she's been writing down two things every single day β€”

  1. The temperature that day (in Β°C)
  2. How many cups she sold

Here's her notebook:

Day Temperature (Β°C) Cups Sold
1 20 8
2 22 10
3 25 14
4 28 18
5 30 22
6 18 5
7 33 27
8 24 12
9 29 20
10 26 15
11 31 24
12 21 9
13 27 17
14 19 6

Today, the forecast says it'll be 32Β°C. Zara stares at her notebook. Somewhere in those numbers is the answer to "how many cups should I make?" β€” but how does she pull it out of a list of numbers?

This is the exact same problem faced by doctors predicting how a patient will respond to a dose of medicine, or companies predicting how much stock to order, the same with scientists predicting how fast a glacier will melt.

Different costumes, same mystery: "Given what I know, what should I expect to happen?"


Building Intuition β€” The Friend-Guessing Game

Here's a game you've possibly played without realizing it.

Imagine being told: "My cousin is 10 years old. How tall do you think she is?"

You don't know her. You've never met her. But you don't shrug and say "no idea" β€” you instantly picture some height. Maybe around 140 cm. Why? Because in your head, without even trying, you've stored a rough pattern from every person you've ever seen:

Younger kids β†’ shorter. Older kids β†’ taller. Adults β†’ tallest (usually). At least that is what your mind's pictorial generally suggesting

You do not have a memorized formula for this. You built it from noticing a pattern across hundreds of people you've seen in your life. That pattern lets you make a reasonable guess about someone you've never met, using just one clue β€” their age.

Zara's temperature notebook is the exact same idea, except her "hundreds of people" are her 14 days of data, and her "clue" is temperature instead of age.


Seeing the Pattern

Let's stop staring at the table and actually look at the data instead.

If you were to plot each day as a dot β€” temperature on the bottom, cups sold on the side β€” you'd get something like this:

graph showing number of cups sold per day and the temperature in centigrade
(picture 14 dots scattered across a graph)

  • Cold days (18–21Β°C) β†’ dots clustered low, around 5–9 cups
  • Warm days (28–31Β°C) β†’ dots clustered high, around 18–24 cups
  • Nothing bounces around randomly β€” hot days are never low-selling days, and cold days are never high-selling days.

Even without doing any math, your eyes can tell something important: as temperature goes up, cups sold goes up too β€” and it does so pretty consistently.

This "shape" β€” a cloud of dots drifting upward together β€” is one of the most important visual patterns in all of data science.

If you see it, you're already halfway to understanding regression.


Naming What We've Found

That upward-drifting relationship between temperature and cups sold has a name:

Trend (informally) β€” its the general direction a pattern moves in.

Data scientists have a more formal word for a predictable, describable relationship like this, but we'll earn that word properly in Phase 2. For now, just sit with this idea:

A pattern is a relationship between two things where knowing one helps you guess the other.

That's it. That's the entire foundation of everything in this series. Every idea coming up β€” regression, best-fit lines, overfitting, regularization β€” is just a more careful, more disciplined way of answering one question:

"Given this pattern, what's my best guess?"


A Tiny Bit of Reasoning (No Scary Math Yet)

Let's not jump to formulas. Let's just reason like Zara would, sitting with her notebook.

She notices:

  • Around 20Β°C, she tends to sell about 7–8 cups.
  • Around 30Β°C, she tends to sell about 20–22 cups.
  • Each time the temperature climbs by about 1Β°C, sales seem to climb by roughly 1 to 1.5 cups.

So for tomorrow's forecast of 32Β°C, Zara can reason:
"That's warmer than my warmest day so far. Sales should be a bit higher than my best day β€” maybe around 26 to 28 cups?"

Notice what just happened. Zara didn't guess randomly. She didn't need a computer. She just used the pattern in her own data to make an educated, reasonable prediction.

Believe it or not, this is exactly what companies pay data scientists large salaries to do β€” only that they work with millions of data points instead of 14.

Plus, they work with tools that are far more precise than "eyeballing datapoints."

Even then, the core idea never changes.


Why This Matters in the Real World

This "notice a pattern, use it to predict" trick shows up everywhere:

  • Weather-based businesses: Ice cream shops, umbrella sellers, and yes β€” lemonade stands β€” all use temperature patterns to plan stock.
  • Ride-sharing apps: Uber and Bolt predict how many drivers will be needed based on time of day and past demand patterns.
  • Hospitals: Predicting how many beds will be needed in the ER based on patterns like day of the week or season.
  • Farmers: Predicting crop yield based on rainfall patterns from past seasons.

In every one of these cases, someone is standing exactly where Zara is standing β€” looking at a notebook full of past numbers, trying to make a smart guess about tomorrow.


πŸ” Cliffhanger

Zara's eyeballing trick worked okay for a rough guess. But here's the problem: if you gave her notebook to five different friends and asked each of them to draw a line through the dots that best represents the pattern, you'd get five different lines.

Which one line is actually most correct?

Is there even such a thing as one "correct" line β€” or is it all just opinion?

(Coming up next: Drawing the "Best Guess" Line.)


πŸ“Œ Key idea to carry forward: A pattern is a relationship where knowing one thing helps you make an educated guess about another.

Everything else we build in this series is just a more precise, more disciplined way of finding and using that pattern.


Phase 2: Drawing the "Best Guess" Line

The Problem

Zara shows her notebook to four friends β€” Malik, Priya, Tom, and Aisha β€” and asks each of them the same thing:

"Can you draw one line through these dots that best shows the pattern? I want to use it to guess tomorrow's sales."

They each grab a ruler and pencil, stare at the scattered dots, and draw:

  • Malik draws a line that passes exactly through the very first and very last dot.
  • Priya draws a line that seems to "hug" the middle of the cloud of dots.
  • Tom draws a line that tilts a bit steeper, trying to match the hottest days more closely.
  • Aisha draws a line that's almost flat, arguing "the sales don't change that much."

Zara looks at four different lines, on the exact same data, drawn by four reasonable people.

Uh oh. If they all "eyeballed" the same dots and got different answers... who's actually right?

Is this just a matter of opinion?
Or is there a correct line hiding in there somewhere?


Building Intuition β€” The Ruler and the Pebbles

Imagine you scatter a handful of pebbles across a table β€” not in a neat row, but roughly following a diagonal path, like they'd been thrown from one corner toward the opposite one.

Now someone hands you a single straight ruler and says: "Lay this ruler down so it represents where the pebbles generally are β€” even though it won't touch most of them."

You wouldn't put the ruler far off to the side, ignoring the pebbles completely. And you probably wouldn't zig-zag it wildly trying to touch every single pebble either (a ruler can't zig-zag anyway β€” it's straight).

Instead, you'd instinctively lay it down somewhere in the "middle" of the scattered pebbles β€” close to as many of them as possible, even if it touches almost none of them exactly.

That instinct β€” finding the straight path that stays closest, on average, to everything β€” is more or less what Zara's friends were trying to do with her dots. Only that they didn't have a precise way to measure "closest, on average."


Seeing the Pattern

Picture Zara's scatter plot again, and now imagine all four lines drawn on top of it at once.

several freehand lines trying to represent best-fit

  • Malik's line (through the first and last dot) actually drifts quite far from several dots in the middle β€” it "represents" the two edge points perfectly but ignores everyone in between.
  • Aisha's near-flat line stays far below the hot-day dots and far above the cold-day dots β€” it barely captures the pattern at all.
  • Priya's and Tom's lines look reasonably close to most dots, but lean in slightly different directions.

Here's the important visual insight: none of these lines touch every single dot β€” and that's actually fine.

A dataset like Zara's isn't a perfectly straight row of points; it's a cloud with a general direction.

Touching every dot is not the goal.

The goal is to find the one straight line that stays as close as possible
=> to ALL the dots,
=> on average,
=> at the same time.


Naming What We've Found

That single straight line β€” the one that best represents the overall trend of a cloud of dots β€” can now be given a proper name:

Regression Line (also known as the Line of Best Fit): a straight line drawn through a set of data points so that it represents the overall relationship between the two variables (in Zara's case, temperature and cups sold) as closely as practically possible.

Or so it goes.

And the process of finding that line β€” of taking messy, scattered real-world data and boiling it down into one clean, usable line β€” is called:

Regression: the process of finding the line (or curve) that best describes the relationship between an input (like temperature) and an output (like cups sold), so that we can use it to make predictions.

Say that word out loud once β€” "Regression" β€” because for the rest of this series, you'll see that everything else builds on top of this one idea.


A Tiny Bit of Math (Just Enough, Nothing Scary)

You've actually seen the formula for a straight line before, probably in a math class, maybe without realizing you'd use it again here:

y = mx + b

Let's translate this into Zara's lemonade world instead of abstract letters:

Cups Sold = (slope) x Temperature + (starting point)

  • Slope (m): How steeply cups sold increases for every extra degree of temperature. A slope of 1.2 means "for every 1Β°C increase, expect about 1.2 more cups sold."
  • Starting point / the intercept (b): Roughly, what sales would look like at with temperatures at 0Β°C (even if Zara never actually sees a day that cold β€” it's just where the line would start if extended that far back). Consider it the baseline.

So if Zara's line ends up being something like:

Cups Sold = 1.2 x Temperature - 16$$

Then for tomorrow's forecast of 32Β°C would be:

Cups Sold = 1.2 x 32 - 16

working out as:

1.2 x 32 = 38.4

38.4 - 16 = 22.4

Zara should expect to sell roughly 22 cups. Notice this isn't a wild guess anymore β€” it's a number that came directly out of the pattern in her own past data.

(We haven't yet explained **how* to workout that exact slope and intercept β€” that's coming very soon, in Phase 3 and 4, once we figure out how to actually measure which line is "best." For now, just understand what the slope and intercept represent.)*


Why This Matters in the Real World

The "regression line" isn't just a lemonade-stand trick β€” it's one of the most widely used tools in the world:

  • Real estate: Predicting a house's price based on its size, using a line fitted to hundreds of past home sales.
  • Fitness apps: Predicting your future running pace based on your training pattern over past weeks.
  • Economics: Predicting how much people will spend based on their income, using decades of national data.
  • Medicine: Predicting how a patient's blood pressure might respond to a given dosage, based on prior patient data.

In every case, someone had a scatter of messy real-world dots, and needed one clean line to make sense of it β€” exactly like Zara.


πŸ” Cliffhanger

Here's the catch. We still haven't actually answered the original question:

Out of Malik's, Priya's, Tom's, and Aisha's four different lines β€” which one is genuinely the "best"?

We've only said "the best line stays closest to the dots on average" β€” but we haven't measured how close each line actually is. Without a way to measure that, "best" is still just an opinion.

We need a scorekeeper. We need a way to turn "this line looks pretty close to the dots" into an actual number we can compare, fairly, between Malik's line and Priya's line.

(Find out next in Phase 3: How Do We Know Which Line Is "Best"?)


πŸ“Œ Key idea to carry forward: A regression line is the single straight line that best represents the overall trend in a cloud of data points β€” and regression is simply the process of finding it, so it can be used to make predictions.


Phase 3: How Do We Know Which Line Is "Best"?

  • Zara's four friends each drew a different regression line through the same dots.
  • We've come to appreciate that regression is the search for the one line that best represents a trend.
  • Cliffhanger: without a way to measure closeness, "best line" is relative.

The Problem

Zara lays out all four lines β€” Malik's, Priya's, Tom's, and Aisha's β€” on top of her scatter plot and studies them side by side.

She can sort of tell that Aisha's flat line looks like it's missing the pattern, and Malik's line seems to swing too wildly. But when she compares Priya's and Tom's lines, she genuinely can't tell which is better just by looking. They both seem "close enough" to the dots.

She needs something better than a gut feeling. She needs a number β€” a score she can calculate for each line, so that whichever line gets the lowest score wins, fair and square, no opinions involved.

The question is: how do you turn "this line looks pretty close to the dots" into an actual measurable number?


Building Intuition β€” The Archery Range

Picture an archery range. Four archers β€” Malik, Priya, Tom, and Aisha β€” each shoot 10 arrows at the same target.

From a distance, all four look "pretty good." Most arrows land somewhere near the bullseye. But the coach doesn't hand out the trophy based on a glance. Instead, for every single arrow, the coach measures exactly how far it landed from the bullseye β€” in centimeters.

  • An arrow that lands dead center: 0 cm off.
  • An arrow that lands a little to the side: maybe 3 cm off.
  • An arrow that landed way off in the grass: 40 cm off.

Once every arrow has a measured distance, the coach can now fairly compare all four archers β€” even if the difference wasn't obvious just by eyeballing the target.

This is exactly the tool Zara needs. Instead of arrows and a bullseye, she has dots and a line. Instead of "distance from the bullseye," she needs "distance from the line."


Seeing the Pattern

Let's zoom in on Priya's line drawn through Zara's dots.

For each day in the notebook, there are two numbers:

  1. What Zara actually sold that day (the dot's real position).
  2. What Priya's line predicts she should have sold, based on that day's temperature (the line's position at that same temperature).

If you draw a small vertical line connecting each dot straight up or down to Priya's line, you get something like a set of little "distance markers"
β€” short for days the line predicted well,
β€” long for days the line predicted poorly.

Now imagine doing that same exercise for Tom's line. Some of his distance markers might be shorter, some might be longer, in different places.

This is the exact visual of the "archery target" idea, just turned sideways. Every dot is an arrow. The line is the bullseye. The vertical distance is how far off the shot landed.


Naming What We've Found

Each of those little vertical distances β€” the gap between what actually happened and what the line predicted β€” has an important name:

Residual (also just called Error): the difference between the actual value and the value the line predicted, for a single data point.

Residual = Actual Value - Predicted Value

For example, say on a day with 28Β°C, Zara actually sold 18 cups, but Tom's line predicted she'd sell 20 cups. Then:

Residual = 18 - 20 = -2

A residual of βˆ’2 means the line overpredicted by 2 cups that day. If the residual had been +2 instead, it would mean the line underpredicted β€” actual sales beat the prediction.

Every single dot on the graph has its own residual.

A "good" line is one where, across ALL the dots, these residuals tend to be small.


A Tiny Bit of Math (Just Enough, Not To Worry)

There's nothing scary here yet β€” it's genuinely just subtraction.

For any data point:

Residual = Actual - Predicted

Let's calculate a few residuals for Tom's line using three days from Zara's notebook (assuming Tom's line predicts roughly 1.1 cups per Β°C above 20Β°C, starting at 5 cups):

Day Temp (Β°C) Actual Cups Tom's Predicted Cups Residual
3 25 14 10.5 +3.5
7 33 27 19.3 +7.7
14 19 6 3.9 +2.1

Notice Tom's line is consistently underpredicting β€” every residual is positive, meaning actual sales are always higher than what his line expects.

That's a useful clue: it hints Tom's line might be sitting a bit too low across the board.

We're not yet combining these residuals into one master score for the whole line (that's Phase 4). For now, just notice that each residual, on its own, tells us how wrong the line was for one specific point.


Why This Matters in the Real World

The idea of "measuring the gap between prediction and reality" is everywhere, not just in lemonade math:

  • Weather forecasting: Meteorologists track the residual between "predicted temperature" and "actual temperature" every single day to improve their models.
  • Medicine: Doctors compare a predicted recovery timeline against a patient's actual recovery to refine treatment models.
  • Manufacturing: Factories measure the residual between a machine's predicted output and its actual output to catch problems early.
  • Sports analytics: Coaches compare a player's predicted performance (based on past stats) against their actual game performance.

In every one of these fields, the residual is the raw ingredient used to judge β€” and eventually improve β€” a model.


πŸ” Cliffhanger

Zara now has a tool to measure how wrong a line is at one single point. But her notebook has 14 days, not 1. Tom's line might do great on some days and terribly on others.

**How do you combine 14 separate residuals into one single, fair, overall score for the whole line.

A score she could compare directly against Priya's, Malik's, and Aisha's lines?**

There's a sneaky trap waiting here too: what happens if you just add all the residuals together, and the positive ones and negative ones cancel each other out β€” making a genuinely bad line look falsely perfect on paper?


πŸ“Œ Key idea to carry forward: A residual is the gap between what actually happened and what a line predicted for one single data point β€” and it's the basic building block we'll use to score, compare, and eventually improve entire regression lines.


Subscribe so as not to miss the next phase...
(To be continued in Phase 4: Turning Mistakes Into a Single Number.)

Top comments (0)