Most pandas tutorials teach you syntax. They show you .mean(), .groupby(), .apply(). But they don't show you how the pieces connect — how a CSV becomes a chart, how a chart becomes a web app, how a web app ends up live on the internet.
I wanted to see that full pipeline for myself. So I built a small project: a student marks dashboard that reads a CSV, calculates results, plots a chart, and runs as a live web app.
The data
Eight students, three subjects:
name,maths,science,english
Arjun,78,85,72
Priya,92,88,95
Rahul,45,50,40
Sneha,88,91,84
Vikram,30,35,28
Anita,67,72,70
Karan,55,60,58
Meera,95,93,97
Small and clean. That's on purpose — the goal was to understand the pipeline, not to wrestle with messy data.
Loading a CSV into a DataFrame
import pandas as pd
df = pd.read_csv("data/marks.csv")
print(df)
Output:
name maths science english
0 Arjun 78 85 72
1 Priya 92 88 95
2 Rahul 45 50 40
3 Sneha 88 91 84
4 Vikram 30 35 28
5 Anita 67 72 70
6 Karan 55 60 58
7 Meera 95 93 97
pd.read_csv() returns a DataFrame — basically a table in memory. Each row has an index (0–7) and each column is a Series.
At this point it's just numbers. No totals, no averages, no conclusions.
Creating new columns
df["total"] = df["maths"] + df["science"] + df["english"]
df["average"] = (df["total"] / 3).round(2)
print(df)
Output:
name maths science english total average
0 Arjun 78 85 72 235 78.33
1 Priya 92 88 95 275 91.67
2 Rahul 45 50 40 135 45.00
3 Sneha 88 91 84 263 87.67
4 Vikram 30 35 28 93 31.00
5 Anita 67 72 70 209 69.67
6 Karan 55 60 58 173 57.67
7 Meera 95 93 97 285 95.00
When you add two columns in pandas, it works row by row automatically. df["maths"] + df["science"] + df["english"] adds the three marks for each student and returns a new Series. Assigning it to df["total"] creates a new column.
.round(2) keeps the average to two decimal places.
Raw marks don't say much on their own. A total and an average turn three numbers into a single metric you can sort and compare.
When the rule defines the result
Define "Pass" as: 40 or more in every subject.
df["result"] = df[["maths", "science", "english"]].apply(
lambda row: "Pass" if all(row >= 40) else "Fail", axis=1
)
print(df[["name", "average", "result"]])
Output:
name average result
0 Arjun 78.33 Pass
1 Priya 91.67 Pass
2 Rahul 45.00 Pass
3 Sneha 87.67 Pass
4 Vikram 31.00 Fail
5 Anita 69.67 Pass
6 Karan 57.67 Pass
7 Meera 95.00 Pass
Line by line:
-
df[["maths", "science", "english"]]selects only the three subject columns. -
.apply(..., axis=1)runs a function on each row.axis=1means row-wise. -
lambda row: ...is a small anonymous function.rowis one student's three marks. -
all(row >= 40)returnsTrueif every mark is 40 or more. - The ternary
"Pass" if ... else "Fail"returns the label.
On this dataset, "every subject ≥ 40" and "average ≥ 40" happen to give the same answer. That's luck. A student with 95, 95, and 20 would pass on average but fail my rule.
The definition of "pass" is a decision. I made it before writing any code, and it shaped every result after.
What the data showed
Looking at the finished table:
- Topper: Meera, average 95.00
- Class average: 69.5
- Only failure: Vikram, average 31.00
- Borderline: Rahul scrapes through at 45.00
Eight rows is small, but the shape of the class is already visible. One clear topper, one clear failure, and a wide middle.
Visualizing with matplotlib
import matplotlib.pyplot as plt
plt.figure(figsize=(8, 5))
plt.bar(df["name"], df["average"], color="steelblue")
plt.title("Average Marks per Student")
plt.xticks(rotation=45)
plt.tight_layout()
plt.savefig("outputs/average_marks.png")
plt.bar() draws the bars, xticks(rotation=45) stops the names from overlapping, and savefig() writes the chart to a file. Meera's bar is the tallest. Vikram's is the shortest. That's obvious at a glance, much faster than reading the table.
Making it usable by anyone
The script works. But only I can run it, only on my laptop, only with Python installed.
Streamlit turns a Python script into a web app. Here's the full app.py:
import streamlit as st
import pandas as pd
st.title("Student Marks Analysis")
df = pd.read_csv("data/marks.csv")
df["total"] = df["maths"] + df["science"] + df["english"]
df["average"] = (df["total"] / 3).round(2)
df["result"] = df[["maths", "science", "english"]].apply(
lambda row: "Pass" if all(row >= 40) else "Fail", axis=1
)
st.dataframe(df)
st.subheader("Average Marks Chart")
st.bar_chart(df.set_index("name")["average"])
-
st.title()adds the heading. -
st.dataframe(df)renders the DataFrame as an interactive table — users can sort columns by clicking headers. -
st.bar_chart(...)renders the chart in the browser.set_index("name")makes student names the X-axis labels.
The pandas logic is identical to the script. Streamlit just converts the output into HTML.
Running it locally:
python -m streamlit run app.py
I had to use python -m streamlit instead of just streamlit. I'd installed it correctly, but Windows didn't know where the executable was, so it kept saying "not recognized." Running it as a Python module sidesteps the PATH issue.
Deploying to the cloud
Streamlit Community Cloud is free and connects directly to GitHub. Sign in, pick the repo, select app.py, click Deploy. It installs dependencies from requirements.txt and gives you a public URL in about two minutes.
One thing went wrong. I'd edited the README on GitHub's website while also editing it locally, so git push was rejected — the remote had changes I didn't have. I fixed it with:
git pull origin main
git push
Git opened Vim to write a merge message. I'd never seen Vim before and had no idea how to get out. Took me a few minutes of searching to find that :wq saves and quits. Small thing, but it's the part I remember most clearly.
Project structure
By the end, the folders looked like this:
student-marks-analysis/
├── data/
│ └── marks.csv
├── outputs/
│ └── average_marks.png
├── src/
│ └── analysis.py
├── app.py
├── requirements.txt
└── README.md
Data in data/, scripts in src/, generated files in outputs/, the web app at the top level. When someone opens the repo, they immediately know what's where.
Dependencies:
pip install pandas matplotlib streamlit
What I'd improve
- Use a real dataset. 8 rows with no missing values, no typos, no duplicates. Real data is messier, and cleaning is where half the work lives.
-
Add filters. A
st.selectbox()for subjects or a slider for marks would make the app actually interactive. - More charts. Subject-wise comparison, or a trend over time if the data had dates.
These aren't flaws. They're the next things to learn.
Summary
The pipeline, end to end:
- Load the CSV with pandas.
- Transform by creating columns and applying rules.
- Visualize so patterns become obvious.
- Ship as a web app and deploy it.
Most tutorials cover steps 1 and 2.
Steps 3 and 4 are where a script becomes something other people can actually use.
It's a small project. But now I get how the pieces connect.

Top comments (0)