DEV Community

Venus-Kennedy
Venus-Kennedy

Posted on

Unsupervised Learning Explained: Finding Hidden Patterns in Data

Machine learning is commonly divided into different learning approaches based on how a model learns from data. One of the most important approaches is unsupervised learning.

In supervised learning, we train a model using data that already has known answers or labels. For example, if we want to predict whether a customer will default on a loan, the training data may contain previous customers and a label showing whether each customer defaulted.

Unsupervised learning is different. The data does not have predefined labels. Instead, the machine learning algorithm examines the data and tries to discover meaningful patterns, structures, relationships, or groups on its own.

This makes unsupervised learning particularly useful when working with large datasets where we do not already know what patterns exist.

What Is Unsupervised Learning?

Unsupervised learning is a machine learning approach where an algorithm learns patterns and structures from data without being given labelled outcomes.

Imagine giving a dataset to a machine learning model containing information about thousands of customers:

  • Age
  • Income
  • Spending
  • Number of purchases
  • Frequency of transactions
  • Account activity

Instead of telling the model which customers belong to which category, we allow the algorithm to identify groups based on similarities in their behaviour.

The model might discover groups such as:

  • Customers who spend frequently
  • Customers who make occasional large purchases
  • Customers with low transaction activity
  • Customers with similar income and spending patterns

These groups were not explicitly provided to the model. The algorithm discovered them from the data.

Supervised vs. Unsupervised Learning

Understanding the difference between supervised and unsupervised learning is essential for beginners.

Feature Supervised Learning Unsupervised Learning
Training data Labelled Unlabelled
Known target Yes No
Main goal Predict outcomes Discover patterns
Common tasks Classification, Regression Clustering, Dimensionality Reduction
Example Predict loan default Group similar customers
Evaluation Often straightforward Can be more challenging

Simple example

Suppose a bank has customer data.

With supervised learning, the bank might ask:

"Can we predict whether this customer will repay their loan?"

With unsupervised learning, the bank might ask:

"Can we discover different types of customers based on their financial behaviour?"

The first problem has a known target.

The second problem involves discovering hidden patterns.

Why Do We Need Unsupervised Learning?

Real-world datasets are often messy and do not always come with labels.

Imagine collecting data from millions of customers. Manually assigning a category to every customer could be expensive and time-consuming.

Unsupervised learning can help organizations explore such data automatically.

Some common uses include:

  • Customer segmentation
  • Fraud detection
  • Market research
  • Recommendation systems
  • Anomaly detection
  • Document analysis
  • Image analysis
  • Data exploration
  • Feature extraction
  • Discovering hidden relationships

It is especially useful during the exploratory stage of a data science project.

The Major Types of Unsupervised Learning

There are several techniques used in unsupervised learning, but three important categories are:

  1. Clustering
  2. Dimensionality reduction
  3. Association rule learning

Let's look at each one.

1. Clustering

Clustering is one of the most common forms of unsupervised learning.

The goal is to divide data points into groups, called clusters, based on similarities.

For example, an online store might have thousands of customers.

Instead of treating all customers the same, clustering could identify groups such as:

  • High-value customers
  • Frequent customers
  • Occasional customers
  • Inactive customers

The business can then design different strategies for each group.

Example

Suppose we have the following customers:

Customer Annual Income Annual Spending
A 50,000 5,000
B 52,000 6,000
C 150,000 80,000
D 145,000 75,000
E 45,000 4,000

A clustering algorithm may discover that A, B and E are similar, while C and D form another group.

The algorithm was not told these groups beforehand.

It discovered them based on the characteristics of the data.

K-Means Clustering

One of the most popular clustering algorithms is K-Means.

The basic idea is to divide observations into a specified number of clusters.

For example:

from sklearn.cluster import KMeans

model = KMeans(n_clusters=3, random_state=42)

model.fit(X)

labels = model.labels_
Enter fullscreen mode Exit fullscreen mode

Here:

  • n_clusters=3 tells the algorithm to create three clusters.
  • fit(X) allows the model to learn from the data.
  • labels contains the cluster assigned to each observation.

The value of K represents the number of clusters we want.

Choosing the right value of K is an important part of the process.

2. Dimensionality Reduction

Datasets can contain hundreds or even thousands of variables.

Working with so many variables can make analysis difficult and computationally expensive.

Dimensionality reduction attempts to represent data using fewer variables while preserving important information.

One popular technique is Principal Component Analysis (PCA).

For example, imagine a dataset with:

100 variables
Enter fullscreen mode Exit fullscreen mode

PCA may help transform the dataset into:

10 principal components
Enter fullscreen mode Exit fullscreen mode

while retaining much of the important variation in the original data.

This can make the dataset easier to:

  • Visualize
  • Analyze
  • Process
  • Model

PCA in Python

A simple example using Scikit-learn is:

from sklearn.decomposition import PCA

pca = PCA(n_components=2)

X_reduced = pca.fit_transform(X)
Enter fullscreen mode Exit fullscreen mode

The resulting data contains two principal components.

This is particularly useful when trying to visualize complex datasets in two dimensions.

3. Association Rule Learning

Association rule learning is another type of unsupervised learning.

It attempts to discover relationships between items or events.

A common example is market basket analysis.

Imagine a supermarket analyzing thousands of transactions.

The data might reveal that customers who purchase:

Bread + Butter
Enter fullscreen mode Exit fullscreen mode

frequently also purchase:

Milk
Enter fullscreen mode Exit fullscreen mode

The business can use these patterns for:

  • Product recommendations
  • Store layout decisions
  • Promotions
  • Cross-selling
  • Online shopping suggestions

One popular algorithm for this type of analysis is the Apriori algorithm.

Anomaly Detection

Unsupervised learning can also help identify unusual observations.

An anomaly is a data point that behaves significantly differently from the majority of the data.

For example, a bank may normally see transactions such as:

KES 500
KES 2,000
KES 5,000
KES 10,000
Enter fullscreen mode Exit fullscreen mode

Suddenly, a transaction of:

KES 900,000
Enter fullscreen mode Exit fullscreen mode

may appear unusual.

This does not automatically mean the transaction is fraudulent. However, it could be flagged for further investigation.

Algorithms such as Isolation Forest can be used for anomaly detection.

from sklearn.ensemble import IsolationForest

model = IsolationForest(random_state=42)

model.fit(X)

predictions = model.predict(X)
Enter fullscreen mode Exit fullscreen mode

Unusual observations can then be investigated further.

How Does Unsupervised Learning Work?

A typical unsupervised learning workflow looks like this:

Raw Data
   ↓
Data Cleaning
   ↓
Exploratory Data Analysis
   ↓
Feature Selection
   ↓
Feature Scaling
   ↓
Choose Algorithm
   ↓
Train Model
   ↓
Discover Patterns
   ↓
Interpret Results
Enter fullscreen mode Exit fullscreen mode

Unlike supervised learning, there is usually no predefined target variable.

The objective is to understand the structure hidden inside the dataset.

The Importance of Data Preparation

Although unsupervised learning does not require labels, the quality of the input data still matters greatly.

Poor-quality data can produce misleading patterns.

Important preprocessing steps may include:

Handling missing values

Missing values can affect algorithms, particularly distance-based methods such as K-Means.

Removing duplicates

Duplicate records can distort the structure of the dataset.

Handling outliers

Extreme values can strongly influence some clustering algorithms.

Scaling features

Suppose we have:

Age: 18–80
Income: 20,000–500,000
Enter fullscreen mode Exit fullscreen mode

Income has a much larger numerical scale than age.

Some algorithms may therefore give income disproportionate influence.

Standardization can help:

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()

X_scaled = scaler.fit_transform(X)
Enter fullscreen mode Exit fullscreen mode

The scaled data can then be used for algorithms such as K-Means.

How Do We Evaluate Unsupervised Learning?

Evaluation can be more difficult than in supervised learning because there may be no known correct answer.

For clustering, one commonly used measure is the Silhouette Score.

The silhouette score measures how well observations fit within their assigned clusters compared with other clusters.

It ranges approximately from:

-1 to +1
Enter fullscreen mode Exit fullscreen mode

A higher score generally indicates better-defined clusters, although the score should not be treated as the only basis for deciding whether a clustering result is useful.

In Python:

from sklearn.metrics import silhouette_score

score = silhouette_score(X, labels)

print(score)
Enter fullscreen mode Exit fullscreen mode

Other approaches include:

  • Examining cluster characteristics
  • Visualizing clusters
  • Comparing different numbers of clusters
  • Using domain knowledge
  • Checking whether the discovered groups make practical sense

Real-World Applications of Unsupervised Learning

Banking and Finance

Unsupervised learning can help identify:

  • Customer segments
  • Unusual transactions
  • Spending patterns
  • Similar financial behaviours
  • Potential areas for further investigation

For example, a bank could group customers according to transaction frequency and account activity.

Marketing

Companies can use clustering to understand customer groups.

Instead of creating one marketing campaign for everyone, organizations can identify groups with similar characteristics and behaviours.

E-Commerce

Online businesses can analyze customer behaviour to identify:

  • Frequently purchased products
  • Customer segments
  • Product relationships
  • Unusual purchasing behaviour

These patterns can support recommendation systems.

Healthcare

Unsupervised learning can be used to explore patient data and identify groups with similar characteristics.

For example, researchers may discover patient groups with similar patterns in medical measurements.

However, clinical applications require careful validation and domain expertise.

Cybersecurity

Unsupervised methods can help identify unusual network behaviour.

For example, if most network activity follows a normal pattern but one device behaves very differently, the activity may be flagged for investigation.

Again, an anomaly is not automatically proof of malicious activity.

A Simple Python Example

Let's create a small clustering example.

import pandas as pd
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler

data = pd.DataFrame({
    "income": [30000, 35000, 40000, 100000, 110000, 120000],
    "spending": [5000, 6000, 7000, 50000, 55000, 60000]
})

scaler = StandardScaler()

X_scaled = scaler.fit_transform(data)

model = KMeans(n_clusters=2, random_state=42, n_init=10)

data["cluster"] = model.fit_predict(X_scaled)

print(data)
Enter fullscreen mode Exit fullscreen mode

The algorithm attempts to divide the customers into two groups based on their income and spending patterns.

The output might look conceptually like:

   income  spending  cluster
0   30000      5000        0
1   35000      6000        0
2   40000      7000        0
3  100000     50000        1
4  110000     55000        1
5  120000     60000        1
Enter fullscreen mode Exit fullscreen mode

The exact cluster labels may vary. Cluster 0 does not inherently mean "low-value" and cluster 1 does not inherently mean "high-value." The numbers are simply identifiers assigned by the algorithm.

Unsupervised Learning in Data Science

Unsupervised learning is especially valuable for exploratory data analysis.

A data scientist may begin with a dataset without knowing exactly what relationships exist.

Instead of immediately building a predictive model, they can use unsupervised techniques to investigate the data.

For example:

Dataset
   ↓
Explore the data
   ↓
Find patterns
   ↓
Identify groups
   ↓
Detect unusual observations
   ↓
Understand important features
   ↓
Build better models
Enter fullscreen mode Exit fullscreen mode

This makes unsupervised learning an important tool in the early stages of many data science projects.

Common Unsupervised Learning Algorithms

Algorithm Main Purpose
K-Means Clustering
Hierarchical Clustering Building nested groups
DBSCAN Density-based clustering and anomaly discovery
PCA Dimensionality reduction
Apriori Association rule mining
Isolation Forest Anomaly detection
Gaussian Mixture Models Probabilistic clustering

Different algorithms are suitable for different types of problems.

There is no single algorithm that works best for every dataset.

Challenges of Unsupervised Learning

Unsupervised learning is powerful, but it also presents several challenges.

1. No predefined answers

Because there are no labels, it can be difficult to determine whether the discovered patterns are meaningful.

2. Choosing the number of clusters

Algorithms such as K-Means require the user to specify the number of clusters.

Choosing an inappropriate number can produce misleading results.

3. Sensitivity to data preparation

Scaling, missing values and outliers can significantly affect the results.

4. Difficult interpretation

A mathematical cluster does not automatically correspond to a meaningful real-world category.

A data scientist must interpret the results using domain knowledge.

5. False patterns

Algorithms can discover patterns that appear interesting but have little practical meaning.

This is why statistical reasoning and business understanding remain important.

Unsupervised Learning vs. Supervised Learning: A Simple Example

Imagine an online shop with customer data.

Supervised learning

The business has historical information showing whether customers purchased a product.

The model learns:

Customer information → Purchase / No Purchase
Enter fullscreen mode Exit fullscreen mode

The goal is to predict future purchases.

Unsupervised learning

The business does not have predefined customer categories.

The model examines:

Customer information → Hidden customer groups
Enter fullscreen mode Exit fullscreen mode

The goal is to discover patterns.

Both approaches are useful, but they answer different questions.

Key Takeaways

  • Unsupervised learning works with unlabelled data.
  • It focuses on discovering hidden structures, relationships and patterns.
  • Clustering groups similar observations together.
  • Dimensionality reduction simplifies datasets with many variables.
  • Association rule learning discovers relationships between items or events.
  • Anomaly detection can identify unusual observations.
  • Data preprocessing remains important even when labels are unavailable.
  • K-Means is one of the most commonly introduced clustering algorithms.
  • PCA is a popular dimensionality reduction technique.
  • Evaluating unsupervised learning can be more challenging than evaluating supervised learning.
  • Domain knowledge is essential when interpreting discovered patterns.

Unsupervised learning gives data scientists a way to explore datasets when predefined answers are unavailable. Instead of learning from labelled examples, algorithms examine the data and attempt to uncover its underlying structure.

From customer segmentation and fraud investigation to market basket analysis and dimensionality reduction, unsupervised learning has many applications across different industries.

For anyone learning data science, understanding unsupervised learning is important because real-world datasets do not always arrive with clear labels or predefined categories. Sometimes, before we can predict something, we first need to understand what patterns are already hiding inside the data.

Top comments (0)