<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mohammed Shahed</title>
    <description>The latest articles on DEV Community by Mohammed Shahed (@shahedr).</description>
    <link>https://dev.to/shahedr</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174435%2F8e07e87f-c768-40b2-b654-5981c847fbc5.jpg</url>
      <title>DEV Community: Mohammed Shahed</title>
      <link>https://dev.to/shahedr</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shahedr"/>
    <language>en</language>
    <item>
      <title>I Turned My Repetitive CSV Checks Into a Python CLI</title>
      <dc:creator>Mohammed Shahed</dc:creator>
      <pubDate>Sat, 10 Oct 2026 02:47:47 +0000</pubDate>
      <link>https://dev.to/shahedr/i-turned-my-repetitive-csv-checks-into-a-python-cli-3c0o</link>
      <guid>https://dev.to/shahedr/i-turned-my-repetitive-csv-checks-into-a-python-cli-3c0o</guid>
      <description>&lt;p&gt;I kept running the same checks whenever I opened a new CSV or Excel file.&lt;/p&gt;

&lt;p&gt;Missing values. Duplicate rows. Mixed numeric and text values. Bad dates. Columns that never change. Numbers that look unusually high or low.&lt;/p&gt;

&lt;p&gt;None of those checks is difficult, but doing them manually every time gets old pretty quickly.&lt;/p&gt;

&lt;p&gt;So I turned that first pass into a small Python command-line tool called &lt;strong&gt;Data Quality Detective&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;dqdetect&lt;/code&gt; profiles a CSV or Excel file and gives you a quick report of the things that probably deserve a closer look before you start deeper analysis.&lt;/p&gt;

&lt;p&gt;The current public release checks for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;dataset shape and column types&lt;/li&gt;
&lt;li&gt;duplicate rows&lt;/li&gt;
&lt;li&gt;missing values by column&lt;/li&gt;
&lt;li&gt;constant columns&lt;/li&gt;
&lt;li&gt;mixed numeric/text values&lt;/li&gt;
&lt;li&gt;invalid values in date-like columns&lt;/li&gt;
&lt;li&gt;IQR-based numeric outliers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One thing I did &lt;strong&gt;not&lt;/strong&gt; want the tool to do was automatically clean the data.&lt;/p&gt;

&lt;p&gt;If a value is missing or looks like an outlier, that does not always mean it is wrong. Sometimes it is actually important. So the tool flags the issue and leaves the decision to the analyst.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install it
&lt;/h2&gt;

&lt;p&gt;The first public release is on PyPI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;dqdetect
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then run it on a file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dqdetect your_file.csv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dqdetect messy_orders.csv
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The command gives you a quick summary in the terminal and creates Markdown and HTML reports.&lt;/p&gt;

&lt;p&gt;Example output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Data Quality Detective
File: messy_orders.csv

Rows: 12
Columns: 8
Duplicate rows: 1

Issues found
- 2 columns contain missing values
- 1 column contains mixed numeric/text values
- 1 date-like column contains invalid values
- 2 numeric columns contain potential IQR outliers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why I built it
&lt;/h2&gt;

&lt;p&gt;The goal is not to replace a full data-validation framework.&lt;/p&gt;

&lt;p&gt;I wanted something lightweight for the point where you have just received a file and want to answer one simple question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What should I inspect before I trust this dataset?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is the scope I want to keep the project focused on.&lt;/p&gt;

&lt;h2&gt;
  
  
  A bug CI caught
&lt;/h2&gt;

&lt;p&gt;One useful part of building this was seeing the value of automated testing in a real project.&lt;/p&gt;

&lt;p&gt;When I first added GitHub Actions, the workflow failed because of a formatting bug in the Markdown report generation.&lt;/p&gt;

&lt;p&gt;I fixed the bug, reran the workflow, and the tests passed.&lt;/p&gt;

&lt;p&gt;It was a small issue, but it made CI feel a lot more practical to me. Instead of manually checking whether every change still works, the project now runs those checks automatically whenever the code changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it is now
&lt;/h2&gt;

&lt;p&gt;The project currently has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a working CLI&lt;/li&gt;
&lt;li&gt;automated tests&lt;/li&gt;
&lt;li&gt;GitHub Actions CI&lt;/li&gt;
&lt;li&gt;a tagged &lt;code&gt;v0.1.0&lt;/code&gt; release&lt;/li&gt;
&lt;li&gt;PyPI publishing&lt;/li&gt;
&lt;li&gt;Markdown and HTML reports&lt;/li&gt;
&lt;li&gt;example data and example output&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I also tested the public install separately using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;dqdetect
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and ran it against the sample dataset to make sure the published package works outside the repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I’m working on next
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;main&lt;/code&gt; branch already has some work for the next version, including machine-readable JSON output.&lt;/p&gt;

&lt;p&gt;A few other things I want to explore:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;configurable thresholds&lt;/li&gt;
&lt;li&gt;schema rules&lt;/li&gt;
&lt;li&gt;PostgreSQL profiling&lt;/li&gt;
&lt;li&gt;comparing two versions of a dataset&lt;/li&gt;
&lt;li&gt;richer HTML reports&lt;/li&gt;
&lt;li&gt;GitHub Action support for automated checks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I want to keep adding things that make the tool more useful without turning it into something unnecessarily complicated.&lt;/p&gt;

&lt;p&gt;If you work with messy datasets, I’d be interested to know what checks you usually run first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/Shahedr/data-quality-detective" rel="noopener noreferrer"&gt;https://github.com/Shahedr/data-quality-detective&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PyPI:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://pypi.org/project/dqdetect/" rel="noopener noreferrer"&gt;https://pypi.org/project/dqdetect/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>opensource</category>
      <category>data</category>
      <category>github</category>
    </item>
  </channel>
</rss>
