DEV Community

Suzanne Orido
Suzanne Orido

Posted on

Pandas, Matplotlib and Seaborn: A Beginner's Guide with Clinic Data

Introduction

My data science learning has now moved on to working with data directly in Python. Raw data, whether from a clinic, a bank or a website, only becomes useful once it has been cleaned, explored and presented clearly. Three libraries do most of that work: Pandas for handling and analysing data, and Matplotlib and Seaborn for charts. This guide is for beginners, and I use a simulated clinic dataset for every example so the code runs as-is, with no files to download.

Setting Up
Install the libraries once from your terminal (or in a Colab or Jupyter cell, put an exclamation mark before each command). Then import them at the top of your notebook:

What Is Pandas, and Why Does It Matter?
Pandas is an open-source library for working with structured data such as CSV files, Excel sheets and database tables. It makes common tasks short and readable: filtering, sorting, calculating, cleaning and summarising.
Its two main data structures are the Series and the DataFrame.
A Series is a single labelled column of values:

A DataFrame is a table of rows and columns, much like a spreadsheet or a patient register:

Creating a Practice Dataset
To avoid needing a file, this code builds 200 simulated clinic visits. It is random data for practice only, not real patient information.

In real work you would load a file instead, with
pd.read_csv("your_file.csv")

Exploring the Data
Start every analysis by looking at what you have:

info() is especially useful. It shows that bmi has fewer entries than the other columns, which tells you values are missing.

Handling Missing Values
Missing values distort results, so check for them first.
You can remove rows with missing values, or fill them in:

The clinical habit of asking why something is missing applies here too. A missing value may mean the measurement was never taken, and that is information in itself. I used the median rather than the mean because it is less affected by extreme values.

Filtering, Sorting and Grouping
Filtering selects the rows you care about, such as patients with a systolic blood pressure of 140 or above:

Sorting puts them in order:

Grouping summarises by category, here is the average blood pressure for each diagnosis:

Why Visualise Data?

People read pictures faster than tables of numbers. A good chart shows trends, comparisons and relationships at a glance. The main chart types and their uses:
• Line charts show trends over time.
• Bar charts compare categories.
• Pie charts show proportions of a whole.
• Histograms show how values are distributed.
• Scatter plots show the relationship between two variables.
• Heatmaps show how strongly several variables relate to each other.

Matplotlib: Full Control
Matplotlib is one of the oldest and most flexible plotting libraries in Python. You can control every detail of a chart, including colours, labels, sizes and grid lines.

Line chart: visits per month

Bar chart: average systolic blood pressure by diagnosis

Pie chart: share of visit types

Seaborn: Faster, Better-Looking Statistical Charts

Seaborn is built on top of Matplotlib and works directly with Pandas DataFrames. It gives you more attractive statistical charts with less code, and most analysts use the two libraries together.

Bar plot: average visit cost by diagnosis

Histogram: distribution of patient ages

Heatmap: correlations between numeric variables

Because this practice data is random, the correlations will be close to zero. With real patient data, a heatmap like this can point you to relationships worth investigating, though correlation alone never proves cause.

Matplotlib or Seaborn?
Use Matplotlib when you need detailed control over how a chart looks. Use Seaborn when you want clean statistical charts quickly. In practice you will often use Seaborn to draw the chart and Matplotlib to adjust the titles, labels and sizes.

Best Practices and Common Mistakes
Best practices
• Clean your data before analysing it.
• Check for missing values and duplicates.
• Choose the chart type that fits the question.
• Always add a title and axis labels, with units where relevant.
• Keep charts simple and readable.

Common mistakes
• Forgetting labels or legends.
• Ignoring missing values.
• Using the wrong chart, such as a pie chart with too many slices.
• Cramming too much into one chart.
• Treating correlation as proof of cause.

Conclusion
Pandas, Matplotlib and Seaborn make up the core toolkit for analysing and presenting data in Python. Pandas cleans, organises and summarises the data, and Matplotlib and Seaborn turn it into charts that other people can understand. For anyone starting in data analytics or data science, these three are well worth learning first. My own advice, coming from medicine, is to practise on data from a field you know, because that is how you notice when a result does not make sense.

Top comments (0)