DEV Community

Cover image for I Turned My Repetitive CSV Checks Into a Python CLI
Mohammed Shahed
Mohammed Shahed

Posted on AI-assisted

I Turned My Repetitive CSV Checks Into a Python CLI

I kept running the same checks whenever I opened a new CSV or Excel file.

Missing values. Duplicate rows. Mixed numeric and text values. Bad dates. Columns that never change. Numbers that look unusually high or low.

None of those checks is difficult, but doing them manually every time gets old pretty quickly.

So I turned that first pass into a small Python command-line tool called Data Quality Detective.

What it does

dqdetect profiles a CSV or Excel file and gives you a quick report of the things that probably deserve a closer look before you start deeper analysis.

The current public release checks for:

  • dataset shape and column types
  • duplicate rows
  • missing values by column
  • constant columns
  • mixed numeric/text values
  • invalid values in date-like columns
  • IQR-based numeric outliers

One thing I did not want the tool to do was automatically clean the data.

If a value is missing or looks like an outlier, that does not always mean it is wrong. Sometimes it is actually important. So the tool flags the issue and leaves the decision to the analyst.

Install it

The first public release is on PyPI:

pip install dqdetect
Enter fullscreen mode Exit fullscreen mode

Then run it on a file:

dqdetect your_file.csv
Enter fullscreen mode Exit fullscreen mode

For example:

dqdetect messy_orders.csv
Enter fullscreen mode Exit fullscreen mode

The command gives you a quick summary in the terminal and creates Markdown and HTML reports.

Example output:

Data Quality Detective
File: messy_orders.csv

Rows: 12
Columns: 8
Duplicate rows: 1

Issues found
- 2 columns contain missing values
- 1 column contains mixed numeric/text values
- 1 date-like column contains invalid values
- 2 numeric columns contain potential IQR outliers
Enter fullscreen mode Exit fullscreen mode

Why I built it

The goal is not to replace a full data-validation framework.

I wanted something lightweight for the point where you have just received a file and want to answer one simple question:

What should I inspect before I trust this dataset?

That is the scope I want to keep the project focused on.

A bug CI caught

One useful part of building this was seeing the value of automated testing in a real project.

When I first added GitHub Actions, the workflow failed because of a formatting bug in the Markdown report generation.

I fixed the bug, reran the workflow, and the tests passed.

It was a small issue, but it made CI feel a lot more practical to me. Instead of manually checking whether every change still works, the project now runs those checks automatically whenever the code changes.

Where it is now

The project currently has:

  • a working CLI
  • automated tests
  • GitHub Actions CI
  • a tagged v0.1.0 release
  • PyPI publishing
  • Markdown and HTML reports
  • example data and example output

I also tested the public install separately using:

pip install dqdetect
Enter fullscreen mode Exit fullscreen mode

and ran it against the sample dataset to make sure the published package works outside the repository.

What I’m working on next

The main branch already has some work for the next version, including machine-readable JSON output.

A few other things I want to explore:

  • configurable thresholds
  • schema rules
  • PostgreSQL profiling
  • comparing two versions of a dataset
  • richer HTML reports
  • GitHub Action support for automated checks

I want to keep adding things that make the tool more useful without turning it into something unnecessarily complicated.

If you work with messy datasets, I’d be interested to know what checks you usually run first.

GitHub:

https://github.com/Shahedr/data-quality-detective

PyPI:

https://pypi.org/project/dqdetect/

Top comments (0)