Descriptive and Inferential Statistics: From Data to Decisions
From Mean and Median to Probability, Hypothesis Testing, Confidence Intervals, Correlation, and Regression
Written by Rakesh Kumar Nayak
Statistics is one of the most important foundations of Data Science, Data Analytics, Machine Learning, and Artificial Intelligence.
Whenever we work with data, two fundamental questions arise:
- What does the data tell us?
- What can we conclude about a larger population from the data?
The first question is mainly answered by Descriptive Statistics, while the second is answered by Inferential Statistics.
In this article, we will go step-by-step from the basics of statistics to probability distributions, sampling, confidence intervals, hypothesis testing, correlation, and regression—with practical examples.
1. What Is Statistics?
Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data.
For example, suppose the salaries of five employees are:
₹25,000, ₹30,000, ₹35,000, ₹40,000, ₹50,000
We may want to know:
- What is the average salary?
- What is the middle salary?
- How much do salaries vary?
- What is the probability that a randomly selected employee earns more than ₹40,000?
- Can we use a sample of employees to estimate the average salary of the entire company?
These questions take us into descriptive statistics, probability, and inferential statistics.
2. Population vs Sample
Before understanding inferential statistics, we need to understand two important concepts.
Population
A population is the complete group we are interested in.
Example:
All employees working in a company.
Sample
A sample is a smaller subset selected from the population.
Example:
100 employees selected from the company.
Studying the entire population is often expensive or time-consuming, so we commonly work with samples.
Example
Population = 10,000 employees
Sample = 500 employees
We calculate statistics from the 500 employees and use those results to learn about the larger population.
3. Descriptive Statistics
Descriptive statistics summarizes and describes the data that we already have.
The major areas include:
- Measures of central tendency
- Measures of dispersion
- Measures of position
- Distribution shape
- Data visualization
4. Measures of Central Tendency
Central tendency tells us where the center of the data lies.
The three most common measures are:
- Mean
- Median
- Mode
5. Mean
The mean is the arithmetic average.
Formula
Mean = Sum of all observations / Number of observations
Suppose five students scored:
60, 70, 80, 90, 100
Mean:
(60 + 70 + 80 + 90 + 100) / 5 = 80
Therefore, the average score is 80.
Python Example
import numpy as np
scores = [60, 70, 80, 90, 100]
print(np.mean(scores))
Output:
80.0
6. Median
The median is the middle value after sorting the data.
Example:
20, 30, 40, 50, 60
Median = 40
For an even number of observations:
20, 30, 40, 50
Median:
(30 + 40) / 2 = 35
Why is Median Important?
The median is less affected by extreme values.
Consider:
20, 25, 30, 35, 500
The mean is strongly affected by 500, while the median remains 30.
Therefore, the median is often useful for highly skewed data such as salaries, house prices, and income.
7. Mode
The mode is the most frequently occurring value.
Example:
10, 20, 20, 30, 40
Mode = 20
Mode is particularly useful for categorical data.
Example:
Red, Blue, Blue, Green, Blue
Mode = Blue
8. Mean vs Median vs Mode
| Measure | Meaning | Common Use |
|---|---|---|
| Mean | Average | Numerical data without strong skew |
| Median | Middle value | Data containing outliers or skew |
| Mode | Most frequent value | Categorical or frequency-based data |
9. Measures of Dispersion
Knowing the average is not enough.
Consider two datasets:
Dataset A:
48, 49, 50, 51, 52
Dataset B:
10, 30, 50, 70, 90
Both have a mean of 50.
However, their variability is very different.
This is why we need measures of dispersion.
Important measures include:
- Range
- Variance
- Standard deviation
- Interquartile range
10. Range
Range measures the difference between the maximum and minimum values.
Formula
Range = Maximum - Minimum
For:
10, 20, 30, 40, 50
Range:
50 - 10 = 40
11. Variance
Variance measures how far observations are spread around the mean.
Population variance:
σ² = Σ(X - μ)² / N
Sample variance:
s² = Σ(X - x̄)² / (n - 1)
A larger variance indicates greater variability.
12. Standard Deviation
Standard deviation is the square root of variance.
σ = √Variance
Standard deviation is one of the most commonly used measures of variability.
Python Example
import numpy as np
data = [10, 20, 30, 40, 50]
print("Mean:", np.mean(data))
print("Variance:", np.var(data))
print("Standard Deviation:", np.std(data))
13. Understanding Distribution Through Visualization
A histogram helps us understand how numerical observations are distributed.
For example, an exam-score dataset can show:
- Center
- Spread
- Skewness
- Possible outliers
- Overall distribution shape
Figure 1: Distribution of Exam Scores — Mean and Median
[Insert Figure 1 here]
14. Quartiles and Percentiles
Quartiles divide ordered data into four parts.
Q1
25th percentile
Q2
50th percentile, which is the median
Q3
75th percentile
The Interquartile Range is:
IQR = Q3 - Q1
IQR is particularly useful for identifying potential outliers.
15. Box Plot
A box plot visually represents:
- Minimum
- Q1
- Median
- Q3
- Maximum
- Potential outliers
Figure 2: Box Plot and Five-Number Summary
[Insert Figure 2 here]
Box plots are especially useful when comparing distributions across multiple groups.
16. Probability
Probability measures the likelihood that an event will occur.
Basic Formula
P(A) = Number of favorable outcomes / Total number of possible outcomes
For a fair coin:
P(Head) = 1 / 2 = 0.5
Therefore, the probability is 50%.
Probability always lies between:
0 ≤ P(A) ≤ 1
Where:
0 = Impossible
1 = Certain
17. Basic Probability Rules
Addition Rule
For mutually exclusive events:
P(A ∪ B) = P(A) + P(B)
Multiplication Rule
For independent events:
P(A ∩ B) = P(A) × P(B)
Complement Rule
P(Aᶜ) = 1 - P(A)
These rules form the foundation of probability calculations.
18. Conditional Probability
Conditional probability measures the probability of an event when another event has already occurred.
Formula
P(A|B) = P(A ∩ B) / P(B)
Example
Suppose we know that a customer purchased a laptop.
What is the probability that the same customer also purchased a laptop bag?
This is an example of conditional probability.
19. Bayes' Theorem
Bayes' theorem allows us to update the probability of an event after receiving new information.
Formula
P(A|B) = [P(B|A) × P(A)] / P(B)
Bayes' theorem is widely used in:
- Spam detection
- Fraud detection
- Medical diagnosis
- Recommendation systems
- Classification
- Naive Bayes algorithms
20. Random Variables
A random variable represents a numerical outcome of a random experiment.
For example:
Let X = number of heads obtained when tossing a coin three times.
Possible values:
0, 1, 2, 3
There are two major types.
Discrete Random Variable
Takes countable values.
Examples:
- Number of customers
- Number of calls
- Number of defective products
Continuous Random Variable
Can take any value within a range.
Examples:
- Height
- Weight
- Temperature
- Time
21. Probability Distributions
A probability distribution describes how probabilities are assigned to possible outcomes.
Important probability distributions include:
- Uniform Distribution
- Bernoulli Distribution
- Binomial Distribution
- Poisson Distribution
- Normal Distribution
- Exponential Distribution
Understanding these distributions is important for Data Science and statistical modeling.
22. Uniform Distribution
In a uniform distribution, outcomes within a specified interval have equal probability density.
For example, a random number generated between 0 and 1 follows a continuous uniform distribution if every value in that interval is equally likely in the density sense.
23. Bernoulli Distribution
Bernoulli distribution represents an experiment with exactly two possible outcomes.
Examples:
- Success / Failure
- Yes / No
- Pass / Fail
- 1 / 0
If:
P(Success) = p
Then:
P(Failure) = 1 - p
Bernoulli distribution is the foundation of the binomial distribution.
24. Binomial Distribution
Binomial distribution models the number of successes in a fixed number of independent Bernoulli trials.
Formula
P(X = k) = C(n,k) × pᵏ × (1-p)ⁿ⁻ᵏ
Example
A coin is tossed 10 times.
What is the probability of getting exactly 5 heads?
Here:
n = 10
p = 0.5
k = 5
Figure 3: Binomial Distribution
[Insert Binomial Distribution Figure here]
25. Poisson Distribution
Poisson distribution is commonly used to model the number of events occurring during a fixed interval of time or space.
Formula
P(X = k) = e⁻λ × λᵏ / k!
Examples:
- Number of customer calls per hour
- Number of website visitors per minute
- Number of machine failures per month
- Number of accidents per day
Figure 4: Poisson Distribution
[Insert Poisson Distribution Figure here]
26. Normal Distribution
The normal distribution is one of the most important distributions in statistics.
It has a characteristic bell-shaped curve.
It is defined by:
μ = Mean
σ = Standard Deviation
The standard normal distribution has:
μ = 0
σ = 1
Figure 5: Standard Normal Distribution
[Insert Normal Distribution Figure here]
27. The 68–95–99.7 Rule
For approximately normally distributed data:
Approximately 68% of observations lie within 1 standard deviation of the mean.
Approximately 95% lie within 2 standard deviations.
Approximately 99.7% lie within 3 standard deviations.
This is called the Empirical Rule.
28. Z-Score
A z-score tells us how many standard deviations an observation is from the mean.
Formula
Z = (X - μ) / σ
Example
Mean = 70
Standard deviation = 10
Student score = 90
Z = (90 - 70) / 10
Z = 2
Therefore, the score is 2 standard deviations above the mean.
29. Inferential Statistics
Now we move from describing data to making conclusions.
Inferential statistics uses sample data to estimate population parameters and test hypotheses.
Major concepts include:
- Sampling
- Sampling distributions
- Central Limit Theorem
- Point estimation
- Confidence intervals
- Hypothesis testing
- p-values
- t-tests
- z-tests
- ANOVA
- Chi-square tests
- Correlation
- Regression
30. Sampling
Suppose a company has 50,000 customers.
Instead of surveying all 50,000 customers, we select 1,000 customers.
The 1,000 customers form our sample.
The objective is to use information from the sample to learn about the larger population.
Good sampling aims to reduce bias and obtain a sample that reasonably represents the target population.
31. Sampling Distribution
Suppose we repeatedly take samples from a population and calculate the mean of every sample.
The collection of those sample means forms the sampling distribution of the sample mean.
Figure 6: Sampling Distribution of the Sample Mean
[Insert Sampling Distribution Figure here]
This concept is fundamental to statistical inference.
32. Central Limit Theorem
The Central Limit Theorem is one of the most important concepts in statistics.
Under common conditions, as the sample size becomes sufficiently large, the sampling distribution of the sample mean approaches a normal distribution, even when the original population is not normally distributed.
This is one reason we can make statistical inferences using sample data.
33. Point Estimation
A point estimate provides one value as an estimate of a population parameter.
For example:
Sample mean = 72
We may use 72 as an estimate of the population mean.
However, a single number does not communicate the uncertainty around the estimate.
This leads to confidence intervals.
34. Confidence Interval
A confidence interval provides a range of plausible values for a population parameter.
A simplified confidence interval for a mean is:
Estimate ± Critical Value × Standard Error
For a 95% confidence interval, the commonly used standard normal critical value is approximately 1.96 when the normal approximation is appropriate.
Example
Suppose:
Sample mean = 72
Standard error = 2
Then:
72 ± 1.96 × 2
The interval is approximately:
68.08 to 75.92
Figure 7: 95% Confidence Interval
[Insert Confidence Interval Figure here]
Important Interpretation
A 95% confidence level does not mean that there is a 95% probability that a particular already-computed interval contains the fixed population parameter.
In frequentist statistics, it means that if we repeatedly constructed intervals using the same method, approximately 95% of those intervals would contain the true parameter.
35. Hypothesis Testing
Hypothesis testing provides a formal framework for evaluating claims about a population.
We generally define two hypotheses.
Null Hypothesis
H₀
The baseline assumption.
Alternative Hypothesis
H₁ or Hₐ
The competing hypothesis.
36. Hypothesis Testing Example
Suppose a company claims:
"Average delivery time is 30 minutes."
We collect a sample of delivery times.
We could formulate:
H₀: μ = 30
H₁: μ ≠ 30
We then select an appropriate statistical test, calculate the test statistic, and evaluate the evidence against the null hypothesis.
37. Significance Level
The significance level is represented by:
α
A commonly used value is:
α = 0.05
It defines the threshold used in the statistical decision rule.
38. P-Value
The p-value measures how compatible the observed data are with the null hypothesis under the assumptions of the statistical test.
A common decision rule is:
If p < α:
Reject H₀.
If p ≥ α:
Do not reject H₀.
Important
A p-value is NOT:
- The probability that H₀ is true
- The probability that the result occurred by chance
- A measure of the size or practical importance of an effect
Statistical significance and practical significance are different concepts.
39. Type I and Type II Errors
Statistical decisions involve uncertainty.
Type I Error
Rejecting a true null hypothesis.
Its probability is commonly associated with:
α
Type II Error
Failing to reject a false null hypothesis.
Its probability is commonly represented by:
β
Statistical power is:
Power = 1 - β
40. Two-Tailed Hypothesis Test
A two-tailed test checks whether a parameter differs from a reference value in either direction.
For α = 0.05 under the standard normal distribution, the critical values are approximately:
-1.96
+1.96
Figure 8: Two-Tailed Hypothesis Test
[Insert Hypothesis Testing Figure here]
41. One-Tailed vs Two-Tailed Tests
One-Tailed Test
Used when the alternative hypothesis is directional.
Example:
H₁: μ > 50
Two-Tailed Test
Used when the alternative hypothesis is non-directional.
Example:
H₁: μ ≠ 50
The choice should be based on the research question and study design, ideally before looking at the results.
42. Common Statistical Tests
| Statistical Test | Typical Application |
|---|---|
| One-Sample t-Test | Compare one sample mean with a reference value |
| Independent t-Test | Compare means of two independent groups |
| Paired t-Test | Compare two measurements from the same subjects |
| Z-Test | Mean/proportion testing under appropriate assumptions |
| ANOVA | Compare means across multiple groups |
| Chi-Square Test | Test association between categorical variables |
| Pearson Correlation | Measure linear association |
| Regression | Model relationships and make predictions |
The correct test depends on the research question, data type, study design, and assumptions.
43. T-Test Example
Suppose we want to investigate whether the average salary of a sample differs from ₹40,000.
We can define:
H₀: μ = ₹40,000
H₁: μ ≠ ₹40,000
A one-sample t-test may be appropriate when the population standard deviation is unknown and the required assumptions are reasonably satisfied.
Python Example
from scipy.stats import ttest_1samp
salary = [38000, 42000, 41000, 39000, 45000]
stat, p_value = ttest_1samp(salary, 40000)
print("Test Statistic:", stat)
print("P-value:", p_value)
44. ANOVA
ANOVA stands for:
Analysis of Variance
It is commonly used to compare the means of three or more groups.
Example:
Suppose we want to compare average salaries across:
- Data Analytics
- Data Science
- Software Engineering
We can formulate:
H₀: μ₁ = μ₂ = μ₃
The alternative hypothesis states that not all population means are equal.
If ANOVA provides sufficient evidence against H₀, further analysis may be required to identify which groups differ.
45. Chi-Square Test
The Chi-square test is commonly used with categorical variables.
Example:
Suppose we want to investigate whether:
Gender
and
Purchase Decision
are associated.
Example data:
| Gender | Purchased | Not Purchased |
|---|---|---|
| Male | 120 | 80 |
| Female | 140 | 60 |
A Chi-square test can be used to assess whether there is statistical evidence of an association between these categorical variables.
46. Correlation
Correlation measures the strength and direction of a linear relationship between two numerical variables.
Pearson correlation coefficient:
-1 ≤ r ≤ +1
Interpretation:
r = +1 → Perfect positive linear relationship
r = 0 → No linear relationship
r = -1 → Perfect negative linear relationship
For example, study hours and exam scores may show a positive correlation.
47. Correlation Does Not Mean Causation
This is one of the most important lessons in statistics.
Suppose ice cream sales increase when swimming pool attendance increases.
The two variables may be correlated.
However, this does not mean ice cream sales cause people to visit swimming pools.
A third variable, such as temperature, could influence both.
Therefore:
Correlation ≠ Causation
48. Regression
Regression is used to model relationships between variables and make predictions under appropriate assumptions.
For simple linear regression:
Y = β₀ + β₁X + ε
Where:
Y = Dependent Variable
X = Independent Variable
β₀ = Intercept
β₁ = Slope
ε = Error Term
Example
We can model exam score based on study hours.
Figure 9: Correlation and Simple Linear Regression
[Insert Correlation & Regression Figure here]
49. Descriptive vs Inferential Statistics
| Descriptive Statistics | Inferential Statistics |
|---|---|
| Describes observed data | Draws conclusions about a population |
| Mean | Confidence Interval |
| Median | Hypothesis Testing |
| Mode | P-Value |
| Range | t-Test |
| Variance | ANOVA |
| Standard Deviation | Chi-Square Test |
| Percentiles | Regression |
| Charts | Population Estimation |
A simple way to remember the difference:
Descriptive Statistics tells us what happened in the data.
Inferential Statistics helps us reason about what the data may imply beyond the observed sample.
50. Complete Statistical Workflow for a Data Science Project
A practical statistics workflow can look like this:
Step 1 — Define the Problem
Understand the business or research question.
Step 2 — Collect Data
Data may come from:
- CSV files
- Excel
- Databases
- APIs
- Web scraping
- Surveys
Step 3 — Clean the Data
Check for:
- Missing values
- Duplicates
- Incorrect data types
- Outliers
- Inconsistent values
Step 4 — Perform Exploratory Data Analysis
Calculate:
- Mean
- Median
- Mode
- Range
- Variance
- Standard deviation
- Quartiles
- Percentiles
Step 5 — Visualize the Data
Use:
- Histograms
- Box plots
- Bar charts
- Scatter plots
- Line charts
- Distribution plots
Step 6 — Understand Probability
Identify appropriate probability concepts and distributions.
Step 7 — Formulate Hypotheses
Define:
H₀
and
H₁
Step 8 — Select the Statistical Test
Choose based on:
- Data type
- Number of groups
- Study design
- Distribution
- Statistical assumptions
Step 9 — Calculate the Test Statistic
Apply the selected statistical method.
Step 10 — Interpret Results
Consider:
- P-value
- Confidence interval
- Effect size
- Practical significance
Step 11 — Communicate the Findings
Translate statistical results into meaningful business or research insights.
51. Practical Data Science Example
Imagine an e-commerce company wants to understand customer spending.
It collects data from 1,000 customers.
Variables include:
- Customer_ID
- Age
- Gender
- Income
- Purchase_Amount
- Products_Purchased
- Rating
Descriptive Analysis
We calculate:
- Average Purchase Amount
- Median Purchase Amount
- Standard Deviation
- Minimum
- Maximum
- Q1
- Q3
Suppose:
Mean Purchase Amount = ₹2,500
Median Purchase Amount = ₹2,200
Standard Deviation = ₹900
These statistics describe the observed sample.
52. Probability Analysis
The company may ask:
"What is the probability that a randomly selected customer spends more than ₹3,000?"
The answer could be estimated using empirical probabilities or an appropriate statistical model, depending on the data and assumptions.
53. Inferential Analysis
Suppose the company wants to compare average spending between two marketing campaigns.
We could formulate:
H₀: μA = μB
H₁: μA ≠ μB
An appropriate statistical test can then be selected based on the study design and assumptions.
The resulting confidence interval, p-value, and effect size can help describe the evidence.
54. Statistics in Machine Learning
Statistics is deeply connected with Machine Learning.
It helps us understand:
Data
Distribution, variability, missing values, and outliers.
Features
Relationships and associations between variables.
Model Evaluation
Performance measurements and uncertainty.
Sampling
Training and testing datasets.
Probability
Classification and probabilistic predictions.
Hypothesis Testing
Evaluating whether observed differences may be explained by random variation.
Regression
Understanding relationships and prediction.
55. Statistics in Generative AI and Modern Data Science
Even with modern Artificial Intelligence and Generative AI, statistical thinking remains important.
Data Scientists may use statistics for:
- Data exploration
- Experiment design
- A/B testing
- Model evaluation
- Sampling
- Uncertainty estimation
- Error analysis
- Feature analysis
- Business decision-making
AI tools can generate code, but understanding the statistical reasoning behind that code remains important.
56. Most Important Statistical Formulas
Mean
x̄ = Σx / n
Population Variance
σ² = Σ(x - μ)² / N
Standard Deviation
σ = √σ²
Z-Score
Z = (X - μ) / σ
Conditional Probability
P(A|B) = P(A ∩ B) / P(B)
Bayes' Theorem
P(A|B) = P(B|A)P(A) / P(B)
Binomial Distribution
P(X=k) = C(n,k)pᵏ(1-p)ⁿ⁻ᵏ
Poisson Distribution
P(X=k) = e⁻λ λᵏ / k!
Confidence Interval
Estimate ± Critical Value × Standard Error
Simple Linear Regression
Y = β₀ + β₁X + ε
57. What Should a Data Scientist Remember?
You do not need to memorize every statistical formula without understanding it.
Focus on understanding:
Descriptive Statistics
- Mean
- Median
- Mode
- Variance
- Standard deviation
- Range
- Quartiles
- Percentiles
- IQR
- Skewness
Probability
- Basic probability
- Conditional probability
- Bayes' theorem
- Independent events
- Random variables
Probability Distributions
- Bernoulli
- Binomial
- Poisson
- Normal
- Uniform
- Exponential
Inferential Statistics
- Sampling
- Sampling distribution
- Central Limit Theorem
- Confidence intervals
- Hypothesis testing
- P-values
- Type I and Type II errors
- Statistical power
Statistical Tests
- t-test
- z-test
- ANOVA
- Chi-square
- Correlation
- Regression
Most importantly, understand when to use a method, why you are using it, what assumptions it requires, and how to interpret the result.
58. Final Takeaway
Statistics is much more than calculating an average.
The complete journey can be remembered as:
Data → Descriptive Statistics → Probability → Sampling → Sampling Distribution → Estimation → Confidence Intervals → Hypothesis Testing → Correlation → Regression → Decision Making
Descriptive statistics helps us understand the data we have.
Probability helps us quantify uncertainty.
Inferential statistics helps us learn from samples and reason about populations.
For anyone preparing for a career in Data Analytics, Data Science, Machine Learning, or AI, statistics provides one of the strongest foundations for working with real-world data.
About the Author
Rakesh Kumar Nayak
Data Science Enthusiast | Data Analytics | Python | SQL | Machine Learning | Power BI | Data Visualization | Generative AI
I am passionate about transforming raw data into meaningful insights and continuously developing my skills in Data Science, Analytics, Machine Learning, and AI.
Written by Rakesh Kumar Nayak
Top comments (0)