DEV Community

Aditya Ranjan
Aditya Ranjan

Posted on

Automating Data Quality Checks in GitHub Actions with DataGuard AI

Automating Data Quality Checks in GitHub Actions with DataGuard AI

Data quality problems rarely wait until a convenient time to appear. A missing customer identifier, an unexpected null value, or a negative transaction amount can move through an ETL pipeline and affect downstream reports before anyone notices.

One practical way to catch these issues earlier is to make data quality checks part of the development workflow.

In this tutorial, I'll show how to use DataGuard AI, an open-source Python data quality and governance toolkit, to scan sample data automatically using GitHub Actions.

The objective is straightforward: every time code changes, run a data quality scan and generate reports developers can inspect.

Why bring data quality into CI/CD?

Software teams routinely automate unit tests, linting, and security checks. Data pipelines benefit from similar practices, especially when datasets and transformation logic evolve.

Automated checks can help teams identify:

  • Missing or invalid values
  • Duplicate identifiers
  • Unexpected data distributions
  • Potentially sensitive information
  • Violations of explicitly configured business rules

These checks do not replace production monitoring, but they can catch problems earlier in the delivery process.

Introducing DataGuard AI

DataGuard AI is an open-source Python toolkit that provides deterministic data quality checks, governance signals, configurable validation rules, and machine-readable reports.

It supports CSV and JSON files, optional Parquet support, DuckDB, and PostgreSQL.

AI-assisted explanations are optional. The core scanning functionality does not require an LLM or an API key.

GitHub: https://github.com/adiranjan25/dataguard-ai

PyPI: https://pypi.org/project/dataguard-ai/

Step 1: Install DataGuard AI

You need Python 3.10 or newer.

Install the public beta:

python -m pip install dataguard-ai==0.2.1
Enter fullscreen mode Exit fullscreen mode

Try the built-in synthetic retail demonstration:

dataguard demo --rows 100
Enter fullscreen mode Exit fullscreen mode

This generates synthetic retail datasets and scans them for data quality and governance risks.

Step 2: Prepare sample data

Create a file called customers.csv:

customer_id,name,email,age
101,Alice,alice@example.com,29
102,Bob,,34
103,Carol,carol@example.com,41
104,David,david@example.com,-5
Enter fullscreen mode Exit fullscreen mode

This intentionally includes a missing email address and an invalid negative age.

Create dataguard.yml:

quality:
  max_null_pct: 5

rules:
  - type: not_null
    column: customer_id
    severity: CRITICAL
  - type: not_null
    column: email
    severity: HIGH
  - type: between
    column: age
    min: 0
    max: 120
    severity: HIGH
Enter fullscreen mode Exit fullscreen mode

These rules express a simple dataset contract: customer identifiers and email addresses must be present, and age must fall within a reasonable range.

Step 3: Run a local scan

Execute:

dataguard scan customers.csv \
  --config dataguard.yml \
  --json-out report.json \
  --html-out report.html
Enter fullscreen mode Exit fullscreen mode

DataGuard AI will inspect the dataset and write two reports:

  • report.json — structured findings suitable for automation
  • report.html — a human-readable report for reviewing findings

Because the sample data deliberately contains invalid values, you should expect findings related to missing email data and the age range rule.

The exact severity, scores, and other findings depend on the tool's configuration and detection behavior.

Step 4: Automate scanning with GitHub Actions

Create this file in your repository:

.github/workflows/data-quality.yml

name: Data Quality Checks

on:
  push:
    branches: [main]
  pull_request:

jobs:
  data-quality:
    runs-on: ubuntu-latest

    steps:
      - name: Checkout repository
        uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.12"

      - name: Install DataGuard AI
        run: python -m pip install dataguard-ai==0.2.1

      - name: Scan customer data
        run: |
          dataguard scan customers.csv \
            --config dataguard.yml \
            --json-out report.json \
            --html-out report.html

      - name: Upload reports
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: dataguard-quality-report
          path: |
            report.json
            report.html
          if-no-files-found: warn
Enter fullscreen mode Exit fullscreen mode

Now, when someone pushes changes to main or opens a pull request, GitHub Actions runs the scanner and uploads the generated reports.

Developers can open the workflow run, find its artifacts, and download the reports for review.

Important limitation: DataGuard AI v0.2.1 generates findings, but the workflow above does not automatically fail merely because a quality issue is detected. The scan step succeeds if execution and report generation succeed. Enforcing a quality gate requires an additional policy step that evaluates the JSON report against agreed thresholds.

Step 5: Review the findings

The HTML report provides an accessible way to inspect quality and governance scores, affected columns, severity, and suggested remediation.

The JSON report is more suitable for downstream automation, such as quality dashboards, policy evaluation, or custom CI checks.

This separation is useful because different teams need different levels of detail: engineers may want structured output, while reviewers may prefer a visual summary.

Where this approach fits

This pattern can be useful for:

  • Data engineering repositories containing sample or test datasets
  • ETL development and transformation validation
  • Data contract experimentation
  • Governance checks during development
  • CI/CD demonstrations and training

For large production datasets, teams should evaluate scanning cost, data sensitivity, execution environment, and performance before adopting the same approach.

Final thoughts

Data quality should not begin only after a pipeline reaches production.

By integrating lightweight, reproducible checks into GitHub Actions, engineering teams can make quality risks visible earlier and establish a foundation for stronger data delivery practices.

DataGuard AI is still a public beta, and community feedback is welcome.

If you work with Python, ETL, data governance, or CI/CD, I'd appreciate your feedback on the installation experience, validation capabilities, and opportunities for improvement.

Explore the source code: https://github.com/adiranjan25/dataguard-ai

Install from PyPI: https://pypi.org/project/dataguard-ai/

Contribute or report an issue: https://github.com/adiranjan25/dataguard-ai/issues

Top comments (0)