Introduction
In my journey of learning Python, one of the things I considered hard a couple of weeks ago was list comprehensions. When I first encountered list comprehensions, I felt intimidated and opted for the classic for loop instead.
However, since I have been challenging myself to do the hard things, I decided it was time to understand them.
In this article, I will share what I have learned and how I have come to understand list comprehension.
The classic for loop
Before looking at list comprehensions, let’s revisit the classic for loop.
As we learnt on Python Loops, for loops are useful when we want to repeat an action for each item in a sequence, such as a list.
Let's start with a simple example.
numbers = [1, 2, 3, 4, 5]
squares = []
for number in numbers:
squares.append(number ** 2)
print(squares)
What is happening in the code above:
- We have a list of numbers.
- We create an empty list called
squares. - The for loop takes each number from numbers and squares it.
- We add the result to squares.
The output:
[1, 4, 9, 16, 25]
We can get the same output using the code below:
numbers = [1, 2, 3, 4, 5]
squares = [number ** 2 for number in numbers]
print(squares)
This is called a list comprehension.
What is a list comprehension?
A list comprehension is a shorter way of creating a new list from an existing iterable.
It allows us to go through each item, perform an operation on it, and collect the results into a new list, all in one line of code.
Going back to the example above, the entire for loop
squares = []
for number in numbers:
squares.append(number ** 2)
is compressed to a single line:
squares = [number ** 2 for number in numbers]
While the second version looks completely different from the classic for loop, it is essentially doing the same thing.
To better understand it, let's break it down in an order that makes sense conceptually;
1. [number ** 2 for number in numbers]
The square brackets indicate that we are creating a new list from the results.
With our original for loop, we had to create an empty list first and then use .append() to add each result to it. With a list comprehension, Python creates the new list for us and places each result into it.
2. for number in numbers
If you have worked with for loops before, this part should look familiar. We are telling Python to go through each item in the numbers list, one at a time, referring to each item as number.
3. number ** 2
For each number, we want to calculate its square. number ** 2 raises the current value of number to the power of 2.
The basic syntax for a list comprehension is:
[expression for item in iterable]
In our example therefore;
[(number ** 2) for (number) in (numbers)]
│ │ |
| | └──Iterable
│ └── Item
└── Expression
List Comprehensions in Data Analysis
Having understood what list comprehensions are and how the syntax works, we can look at some practical ways they can be used in data analysis.
List Comprehensions in Data Cleaning
List comprehensions are useful when you need to apply the same simple cleaning operation to every item in a list.
For example, if a health access dataset has column names with inconsistent spacing and capitalization:
columns = [
"Patient ID",
"AGE ",
" Weight",
"Health Facility",
" County",
"Insurance status"
]
clean_columns = [column.strip().lower() for column in columns]
print(clean_columns)
Output:
['patient id', 'age', 'weight', 'health facility', 'county', 'insurance status']
-
.strip()removes unnecessary spaces, while.lower()converts the column names to lowercase.
List Comprehensions in Filtering Data and EDA
List comprehensions can also be useful during exploratory data analysis when you need to filter data based on a condition.
For example, to identify respondents who live more than 10 km from a health facility:
respondents = [
{"id": "HH001", "distance": 5},
{"id": "HH002", "distance": 12},
{"id": "HH003", "distance": 8},
{"id": "HH004", "distance": 15},
{"id": "HH005", "distance": 21}
]
long_distance = [respondent["id"] for respondent in respondents
if respondent["distance"] > 10
]
print(long_distance)
Output:
['HH002', 'HH004', 'HH005']
- The
ifcondition filters the respondents, ensuring that only those who live more than 10 km from a health facility are included in the new list.
List Comprehensions in Statistical Analysis
List comprehensions can also be useful in statistical analysis when working with several groups. For example, when performing tests such as ANOVA or Kruskal-Wallis, you may need to separate your data into groups before passing them to the test.
A list comprehension provides a cleaner way to create these groups without writing repetitive code for each one.
groups = [
group["bill_amount_ksh"].dropna() for _, group in health_data.groupby("department")
]
- Here, the comprehension goes through each department, removes missing bill amounts, and creates a list containing the bill amounts for each group. The resulting groups can then be passed to the statistical test.
List comprehensions can also be used when applying the same statistical test to several groups. For example, the Shapiro-Wilk test can be used to assess the normality of bill amounts within each department.
normality_results = [
shapiro(group["bill_amount_ksh"].dropna()) for _, group in health_data.groupby("department")
]
In this example, the Shapiro-Wilk test is applied to the bill amounts in each department, and the results are collected into a list. This can help you assess whether the groups meet the normality assumption for certain statistical tests.
When Not to Use List Comprehensions
While list comprehensions provide a more compact way of writing code, there are situations where a classic for loop is a better option:
When the operation is complex. If you need to perform several steps, use multiple conditions, or the comprehension becomes long and difficult to follow, a regular
forloop is usually clearer.When you need more control over the loop. If you need to use
break,continue, or handle errors withtryandexcept, a regularforloop gives you more flexibility and makes the logic easier to follow.When you are not creating a new list. List comprehensions are designed to create lists. If you are performing an action such as printing values, writing to a file, or modifying existing data, a regular
forloop is generally more appropriate.
The aim of using a list comprehension is simplicity and clarity, not simply writing fewer lines of code. If those are not achieved because the logic is too complex or difficult to follow, a regular for loop is the better option.
Conclusion
Learning list comprehensions has reminded me that sometimes a concept can look more complicated than it actually is. What initially seemed unfamiliar and intimidating became much easier once I understood the pattern behind it.
Like most things in Python, becoming comfortable with list comprehensions comes with constant practice. The more you use them in real problems, the easier it becomes to recognise when they are useful and write them without having to think too much about the syntax.
Top comments (0)