This week in my data science class at LuxDev. HQ, we explored various ways in which Pandas is essential for examining our data. Pandas is a Python package with various analytical tools built into it. Pandas introduces an amazing object called a DataFrame, which is like a 2-dimensional list.
Some of the examining techniques we looked at include
1. Viewing your dataframe using .head(n) and .tail(n)
Here n is the number of rows you want to see. Dot head, as the name suggests, is used to view the head or top rows of our dataframe; dot tail is used to view rows at the end of our dataframe
df.tail
2. Checking for null values
Here we used .isnull and .isna(), but you get boolean results, so we add . sum to make it easier to interpret
df.isnull
3. Getting summary statistics
Here we used describe(), which prints out the summary statistics for the integer columns in our dataframe.
4. Using locators in a dataframe
We also explored how loc and iloc can be utilized. iloc is an integer locator and is used to track specific rows/columns using their indexes (integer locations), while loc uses index for rows and actual column names for columns.
Apart from these, Pandas has many other functions that make data screening super easy, and it's a must-have tool for any data scientist.

Top comments (0)