I kept running the same checks whenever I opened a new CSV or Excel file.
Missing values. Duplicate rows. Mixed numeric and text values. Bad dates. Columns that never change. Numbers that look unusually high or low.
None of those checks is difficult, but doing them manually every time gets old pretty quickly.
So I turned that first pass into a small Python command-line tool called Data Quality Detective.
What it does
dqdetect profiles a CSV or Excel file and gives you a quick report of the things that probably deserve a closer look before you start deeper analysis.
The current public release checks for:
- dataset shape and column types
- duplicate rows
- missing values by column
- constant columns
- mixed numeric/text values
- invalid values in date-like columns
- IQR-based numeric outliers
One thing I did not want the tool to do was automatically clean the data.
If a value is missing or looks like an outlier, that does not always mean it is wrong. Sometimes it is actually important. So the tool flags the issue and leaves the decision to the analyst.
Install it
The first public release is on PyPI:
pip install dqdetect
Then run it on a file:
dqdetect your_file.csv
For example:
dqdetect messy_orders.csv
The command gives you a quick summary in the terminal and creates Markdown and HTML reports.
Example output:
Data Quality Detective
File: messy_orders.csv
Rows: 12
Columns: 8
Duplicate rows: 1
Issues found
- 2 columns contain missing values
- 1 column contains mixed numeric/text values
- 1 date-like column contains invalid values
- 2 numeric columns contain potential IQR outliers
Why I built it
The goal is not to replace a full data-validation framework.
I wanted something lightweight for the point where you have just received a file and want to answer one simple question:
What should I inspect before I trust this dataset?
That is the scope I want to keep the project focused on.
A bug CI caught
One useful part of building this was seeing the value of automated testing in a real project.
When I first added GitHub Actions, the workflow failed because of a formatting bug in the Markdown report generation.
I fixed the bug, reran the workflow, and the tests passed.
It was a small issue, but it made CI feel a lot more practical to me. Instead of manually checking whether every change still works, the project now runs those checks automatically whenever the code changes.
Where it is now
The project currently has:
- a working CLI
- automated tests
- GitHub Actions CI
- a tagged
v0.1.0release - PyPI publishing
- Markdown and HTML reports
- example data and example output
I also tested the public install separately using:
pip install dqdetect
and ran it against the sample dataset to make sure the published package works outside the repository.
What Iām working on next
The main branch already has some work for the next version, including machine-readable JSON output.
A few other things I want to explore:
- configurable thresholds
- schema rules
- PostgreSQL profiling
- comparing two versions of a dataset
- richer HTML reports
- GitHub Action support for automated checks
I want to keep adding things that make the tool more useful without turning it into something unnecessarily complicated.
If you work with messy datasets, Iād be interested to know what checks you usually run first.
Top comments (0)