DEV Community

Shank
Shank

Posted on AI-assisted

The Pandas Loc vs iLoc confusion

The other day when I was going through a course on pandas to brush up my knowledge, I came across loc and iloc properties of pandas DataFrame, two ways to pull rows and columns out of a DataFrame. Both the properties do the similar job, the answer you get is a row or column or a slice of a DataFrame you want from a pandas DataFrame or Series. But little did I know how confusing it got for people and how I fell in the same trap myself (many times!).

What are iloc and loc?

Before we start talking about this loc and iloc, we will first learn what both of them are. We will be assuming you already have a basic understanding of Python, what Pandas is and what does Dataframe and Series mean.

To search for a given data ( be it row, column, or a single data point) in a Dataframe, Pandas gives 2 useful properties that come with the Dataframe you create. They are .loc and .iloc.

Both iloc and loc can help you to retrieve a row or column or a combination of row and column with help of slicing or can give you a single data point. The difference between them is loc returns based on the label of the index you give it, while .iloc gives you the data based on the position itself. But what if you don’t assign a label to your Dataframe? when you do .loc[5] and .iloc[5] you get the same result, but they are different, how?

This was one of the 2 confusing steps I was stuck on when I was learning Pandas and still even the most experienced professional can fall into this trap.

Confusion point 1: Default Labels

Let’s say you have imported a DataFrame, in this example I will import the Video Game Sales dataset from kaggle.

#Import pandas 
import pandas as pd

#read the file
df = pd.read_csv(file_path)
df.head()
Enter fullscreen mode Exit fullscreen mode

Importing Videogame Sales Dataset

As we can see, it shows the video game sales of different games across North America, Europe, Japan and others. They also have Global total sales.

We will now look at what will happen if we try to do .loc and .iloc with number 5.

df.loc[5]
Enter fullscreen mode Exit fullscreen mode

.loc property

df.iloc[5]
Enter fullscreen mode Exit fullscreen mode

.iloc property

You can see that when we ran it, we received the 6th row of the DataFrame (remember index starts from 0).
And we haven’t set a label name yet. Then that means by default there are no labels, right? Wrong.
Let’s see what happens when I do the same command after I sort the Dataframe by Global Sales.

loc and iloc results after sorting

What’s this? We have 2 different results here, what's going on?

We have loc returning what we had previously but then we have iloc returning a new row. This is the point of confusion, many assume when no labels are given, loc and iloc behave similarly. But in reality, there is a label by default and it is set to be the position itself.
When you make a DataFrame by default the label starts from 0 and ends with n-1, where n is the total number of rows in the DataFrame. This is the same as the position. Which is why when we ran df.loc[5] and df.iloc[5], we got position 5.

When we sorted the DataFrame according to the Global sales column, the record shifts from its original position to its new position according to the ascending order of Global sales. This is why when you do df.iloc[5] we get a different value since the position 5 is now occupied by the game with 6th lowest Global sales. While the original data in 5th position is now elsewhere. Now when we do df.loc[5] however, we are now looking at the label of “5” and not the position 5. The label is still tied to what the original row at 5th position was, and thus you get the same row back before we sorted the DataFrame.

This point of confusion is common to have from newbies to professionals as well.
The next point of confusion we are going to look at is extremely common for new learners of Pandas.

Confusion point 2: Slicing

Let’s go back to our original DataFrame before we sorted it. Let’s say I want to retrieve rows from 2 to 6.
Since we already know before sorting and by default .loc and .iloc returns the same result because label being the same as position, we shall use both to see what happens.

df.loc[2:6]
Enter fullscreen mode Exit fullscreen mode

slicing DataFrame with .loc

df.iloc[2:6]
Enter fullscreen mode Exit fullscreen mode

slicing DataFrame with .iloc

Wait a minute, why is df.loc returning more values than df.iloc?
Now this is where loc and iloc differ and it’s important new learners know about it.

As you can see .loc gives the label of 6 while .iloc doesn't return the row at position 6. This is because slicing in iloc and loc behaves differently.

.loc includes the final element in the range you are specifying, so if you ask for rows from range 2 to 6, it returns the rows with label 2, 3, 4, 5 and 6. While if you do .iloc, it excludes the final element in the range and thus it would return 2, 3, 4, and 5th position, while ignoring position 6.

Why does .loc include the end point?

There is actually a reasonable explanation to this. We know iloc follows python’s standard way of how range is handled in slicing. But in case of .loc, the range is based on labels, and labels can be anything from strings, dates, or unordered numbers, so there’s no reliable “next label” to stop before. When we do df.loc[‘a’:’d’] for example, excluding ‘d’ would mean you’d have to know which label comes after it and which is why the .loc property doesn’t exclude the endpoint.

What will be returned when you run df.loc[2:6] after sorting?

Experiment and run this on your dataset of choice, and see what answer you get. Does the number of rows returned match the rows you get when you do this before you sort it?

This slicing mechanic is important to keep note of because in standard python we know slicing commands will always exclude the last element. So it is completely normal for someone to get confused and assume it might work the same way when doing the .loc property of DataFrame. Having a mental note of this will save tons of time in debugging why a loc property is not behaving properly.

Takeaways:

So to wrap up the two points of confusion we looked at. First, a DataFrame always has labels, even when you don’t set any. By default they happen to match the positions, which is why loc and iloc seem to give identical results until you sort or filter your data, the labels stay connected to their original rows while the positions change. Second, loc slices include the endpoint while the iloc slices don’t, because labels don’t have the reliable “next” value to stop before.

If you ever want the loc and iloc to line up again, df.reset_index(drop=True) gives you a fresh set of labels that match their current positions.

I fell into both of these traps more times than I’d like to admit whenever I worked with pandas. Hopefully, the next time the loc property returns a row you didn’t expect, or one extra row you didn’t ask for, you’ll know exactly where to look first.


Answer to the Exercise:
With the video game dataset, running df.loc[2:6] in the sorted array would give you an empty DataFrame, this is because loc slices from wherever label 2 sits to wherever label 6 sits in the current order.

Top comments (0)