Automating Data Quality Checks in GitHub Actions with DataGuard AI
Data quality problems rarely wait until a convenient time to appear. A missing customer identifier, an unexpected null value, or a negative transaction amount can move through an ETL pipeline and affect downstream reports before anyone notices.
One practical way to catch these issues earlier is to make data quality checks part of the development workflow.
In this tutorial, I'll show how to use DataGuard AI, an open-source Python data quality and governance toolkit, to scan sample data automatically using GitHub Actions.
The objective is straightforward: every time code changes, run a data quality scan and generate reports developers can inspect.
Why bring data quality into CI/CD?
Software teams routinely automate unit tests, linting, and security checks. Data pipelines benefit from similar practices, especially when datasets and transformation logic evolve.
Automated checks can help teams identify:
- Missing or invalid values
- Duplicate identifiers
- Unexpected data distributions
- Potentially sensitive information
- Violations of explicitly configured business rules
These checks do not replace production monitoring, but they can catch problems earlier in the delivery process.
Introducing DataGuard AI
DataGuard AI is an open-source Python toolkit that provides deterministic data quality checks, governance signals, configurable validation rules, and machine-readable reports.
It supports CSV and JSON files, optional Parquet support, DuckDB, and PostgreSQL.
AI-assisted explanations are optional. The core scanning functionality does not require an LLM or an API key.
GitHub: https://github.com/adiranjan25/dataguard-ai
PyPI: https://pypi.org/project/dataguard-ai/
Step 1: Install DataGuard AI
You need Python 3.10 or newer.
Install the public beta:
python -m pip install dataguard-ai==0.2.1
Try the built-in synthetic retail demonstration:
dataguard demo --rows 100
This generates synthetic retail datasets and scans them for data quality and governance risks.
Step 2: Prepare sample data
Create a file called customers.csv:
customer_id,name,email,age
101,Alice,alice@example.com,29
102,Bob,,34
103,Carol,carol@example.com,41
104,David,david@example.com,-5
This intentionally includes a missing email address and an invalid negative age.
Create dataguard.yml:
quality:
max_null_pct: 5
rules:
- type: not_null
column: customer_id
severity: CRITICAL
- type: not_null
column: email
severity: HIGH
- type: between
column: age
min: 0
max: 120
severity: HIGH
These rules express a simple dataset contract: customer identifiers and email addresses must be present, and age must fall within a reasonable range.
Step 3: Run a local scan
Execute:
dataguard scan customers.csv \
--config dataguard.yml \
--json-out report.json \
--html-out report.html
DataGuard AI will inspect the dataset and write two reports:
-
report.json— structured findings suitable for automation -
report.html— a human-readable report for reviewing findings
Because the sample data deliberately contains invalid values, you should expect findings related to missing email data and the age range rule.
The exact severity, scores, and other findings depend on the tool's configuration and detection behavior.
Step 4: Automate scanning with GitHub Actions
Create this file in your repository:
.github/workflows/data-quality.yml
name: Data Quality Checks
on:
push:
branches: [main]
pull_request:
jobs:
data-quality:
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install DataGuard AI
run: python -m pip install dataguard-ai==0.2.1
- name: Scan customer data
run: |
dataguard scan customers.csv \
--config dataguard.yml \
--json-out report.json \
--html-out report.html
- name: Upload reports
if: always()
uses: actions/upload-artifact@v4
with:
name: dataguard-quality-report
path: |
report.json
report.html
if-no-files-found: warn
Now, when someone pushes changes to main or opens a pull request, GitHub Actions runs the scanner and uploads the generated reports.
Developers can open the workflow run, find its artifacts, and download the reports for review.
Important limitation: DataGuard AI v0.2.1 generates findings, but the workflow above does not automatically fail merely because a quality issue is detected. The scan step succeeds if execution and report generation succeed. Enforcing a quality gate requires an additional policy step that evaluates the JSON report against agreed thresholds.
Step 5: Review the findings
The HTML report provides an accessible way to inspect quality and governance scores, affected columns, severity, and suggested remediation.
The JSON report is more suitable for downstream automation, such as quality dashboards, policy evaluation, or custom CI checks.
This separation is useful because different teams need different levels of detail: engineers may want structured output, while reviewers may prefer a visual summary.
Where this approach fits
This pattern can be useful for:
- Data engineering repositories containing sample or test datasets
- ETL development and transformation validation
- Data contract experimentation
- Governance checks during development
- CI/CD demonstrations and training
For large production datasets, teams should evaluate scanning cost, data sensitivity, execution environment, and performance before adopting the same approach.
Final thoughts
Data quality should not begin only after a pipeline reaches production.
By integrating lightweight, reproducible checks into GitHub Actions, engineering teams can make quality risks visible earlier and establish a foundation for stronger data delivery practices.
DataGuard AI is still a public beta, and community feedback is welcome.
If you work with Python, ETL, data governance, or CI/CD, I'd appreciate your feedback on the installation experience, validation capabilities, and opportunities for improvement.
Explore the source code: https://github.com/adiranjan25/dataguard-ai
Install from PyPI: https://pypi.org/project/dataguard-ai/
Contribute or report an issue: https://github.com/adiranjan25/dataguard-ai/issues
Top comments (0)