Introduction
Many people hear the word statistics and think of dense formulas and confusing charts. As a doctor moving into data science, I have come to see it differently. Programming languages and machine learning get the attention, but statistics is what makes the results trustworthy. It tells you whether a pattern is real, how sure you can be, and what to do next. To show this, I will walk through a made-up but realistic example from a clinic. The numbers are illustrative, not from a real study.
The Scenario
A busy outpatient clinic notices that many patients with hypertension are missing their follow-up appointments. Missed follow-ups matter: uncontrolled blood pressure raises the risk of stroke and kidney disease, and every empty slot is a wasted opportunity for care. The clinic's analyst is asked a simple question: who is missing appointments, and what can we do about it?
The data available includes age, diagnosis, number of previous visits, distance from the clinic, whether the patient received a reminder, and whether they attended. Thousands of rows, no obvious story. Statistics is how the story is found.
Step 1: Describe What Is Happening
The first step is descriptive statistics: averages, counts and rates that summarise the data. Instead of reading thousands of rows, the analyst asks how the missed-appointment rate differs across groups.
The output shows that 20% of patients under 40 missed appointments, 30% of those aged 40 to 64, and 50% of those aged 65 and above. A pattern that was invisible in the raw rows is now clear: older patients are the most likely to miss follow-up.
Step 2: Estimate Risk
Next comes probability. The question changes from "what happened?" to "how likely is it to happen again?" By comparing groups, the analyst can estimate the chance that a patient will miss an appointment, depending on their characteristics. Patients who live far away, who have missed a visit before, or who never received a reminder may each carry a higher probability. Knowing this lets the clinic focus its limited resources on the patients who need the most help.
Step 3: Test the Idea Properly
Suppose the clinic believes that sending SMS reminders will improve attendance. It tries them with one group and compares them with a group that gets none. Imagine 60 of 100 patients attend without reminders, and 78 of 100 attend with reminders. That looks like an improvement, but could it just be chance?
This is what hypothesis testing answers.
A chi-square test checks whether the difference is larger than random variation would plausibly produce.
The p-value comes out well below 0.05, so the difference is unlikely to be due to chance alone. The clinic now has evidence, not just an impression, to justify investing in reminders. In medicine we make the same distinction between an anecdote and a trial result.
Step 4: Predict Who Is at Risk
Statistics also underlies machine learning.
A model such as logistic regression can use many factors at once (age, distance, past attendance, reminders) and produce a risk score for each patient. The clinic can then call high-risk patients before their appointment, offer a telemedicine visit instead, or arrange transport.
The model is not magic. It is statistics applied at scale, and its quality depends on the quality of the data and the care taken in checking it.
Why Statistics Matters in Data Science
• It summarises large datasets into something a human can read.
• It reveals patterns that raw rows hide.
• It measures uncertainty, so you know how far to trust a result.
• It tests ideas before money and effort are committed.
• It supports predictions that lead to earlier, better decisions.
The same logic applies in banking, retail, education and technology. A model without sound statistical thinking can look impressive and still be wrong.
Conclusion
Coding tools and AI models get most of the attention in data science, but they sit on top of statistics. In the clinic example, statistics found who was missing appointments, estimated their risk, showed that reminders actually worked, and powered a model to prevent missed care. For a clinician, none of this is foreign. We already weigh evidence, think in probabilities and ask whether a result could be chance. Statistics gives those habits a formal language, and learning it is one of the most valuable steps toward working with health data.


Top comments (0)