DEV Community

Cover image for Stop bad data at the pull request with Great Expectations and GitHub Actions
Laura Chicovis
Laura Chicovis

Posted on

Stop bad data at the pull request with Great Expectations and GitHub Actions

By the end of this post, a pull request that breaks your data will fail in GitHub before anyone can merge it.

Most data quality checks run after the data lands. But often, what broke it was a code change approved an hour earlier. So let's check it right there, in the pull request, with three pieces:

  1. A data contract in YAML.
  2. A Python script that turns it into Great Expectations (GX) checks.
  3. A GitHub Actions workflow that runs the script on every pull request.

All the files are below, ready to copy.

What you need

Python 3.12, a GitHub repository and this requirements.txt:

great_expectations==1.23.2
pandas==2.3.3
pyyaml==6.0.3
Enter fullscreen mode Exit fullscreen mode

Note that this uses the GX 1.x API, which is quite different from the older 0.x.

Step 1. The pipeline

Raw orders arrive with messy status values, and a small function cleans them up.

data/raw_orders.csv:

order_id,customer_id,status,quantity,unit_price,currency,created_at
1001,C-001,Shipped,2,19.90,USD,2026-09-01
1002,C-002,delivered,1,249.00,USD,2026-09-01
1003,C-003,CANCELLED,3,9.50,USD,2026-09-02
1004,C-001,pending,1,79.00,USD,2026-09-02
1005,C-004,Delivered,5,4.99,USD,2026-09-03
1006,C-005,shipped,2,120.00,USD,2026-09-03
1007,C-002,pending,1,15.00,USD,2026-09-04
1008,C-006,delivered,4,32.50,USD,2026-09-04
Enter fullscreen mode Exit fullscreen mode

pipeline/build_orders.py:

import pandas as pd


def build_orders(raw: pd.DataFrame) -> pd.DataFrame:
    """Turn raw order events into the clean `orders` table."""
    orders = raw.copy()
    orders["status"] = orders["status"].str.strip().str.lower()
    orders = orders[orders["status"] != "cancelled"]
    orders["total_amount"] = (orders["quantity"] * orders["unit_price"]).round(2)
    return orders[
        ["order_id", "customer_id", "status", "total_amount", "currency", "created_at"]
    ]
Enter fullscreen mode Exit fullscreen mode

In CI, we don't test production data. Instead, we test what the code does to a known input, so the check stays fast and predictable.

Step 2. The contract

contracts/orders.yml:

# Data contract for the `orders` table.
# Changing this file changes what downstream teams can rely on.
table: orders
owner: data-platform@yourcompany.com
columns:
  order_id:
    required: true
    unique: true
  customer_id:
    required: true
    pattern: "^C-\\d{3}$"
  status:
    required: true
    allowed_values: [pending, shipped, delivered]
  total_amount:
    required: true
    min: 0.01
    max: 10000
  currency:
    required: true
    allowed_values: [USD]
  created_at:
    required: true
Enter fullscreen mode Exit fullscreen mode

Why YAML? Because anyone on the team can read it, and any change to a rule shows up as a diff that needs approval.

Step 3. From contract to checks

The first half of validate_contract.py turns each rule into a GX Expectation:

import sys

import great_expectations as gx
import pandas as pd
import yaml

from pipeline.build_orders import build_orders

E = gx.expectations


def expectations_from_contract(contract: dict) -> list:
    """Translate each column rule in the YAML contract into a GX Expectation."""
    columns = contract["columns"]
    exps = [E.ExpectTableColumnsToMatchSet(column_set=list(columns), exact_match=True)]

    for name, rules in columns.items():
        if rules.get("required"):
            exps.append(E.ExpectColumnValuesToNotBeNull(column=name))
        if rules.get("unique"):
            exps.append(E.ExpectColumnValuesToBeUnique(column=name))
        if "allowed_values" in rules:
            exps.append(E.ExpectColumnValuesToBeInSet(
                column=name, value_set=rules["allowed_values"]))
        if "pattern" in rules:
            exps.append(E.ExpectColumnValuesToMatchRegex(
                column=name, regex=rules["pattern"]))
        if "min" in rules or "max" in rules:
            exps.append(E.ExpectColumnValuesToBeBetween(
                column=name, min_value=rules.get("min"), max_value=rules.get("max")))
    return exps
Enter fullscreen mode Exit fullscreen mode

The ExpectTableColumnsToMatchSet check matters most. With exact_match=True, any renamed, dropped or extra column fails the contract.

Step 4. Run and exit

The second half goes in the same file, right below:

def main() -> int:
    with open("contracts/orders.yml") as f:
        contract = yaml.safe_load(f)

    orders = build_orders(pd.read_csv("data/raw_orders.csv"))

    context = gx.get_context(mode="ephemeral")
    batch_def = (
        context.data_sources.add_pandas("pipeline_output")
        .add_dataframe_asset(contract["table"])
        .add_batch_definition_whole_dataframe("ci_run")
    )
    suite = context.suites.add(gx.ExpectationSuite(name=f"{contract['table']}_contract"))
    for exp in expectations_from_contract(contract):
        suite.add_expectation(exp)

    validation = context.validation_definitions.add(
        gx.ValidationDefinition(name="contract_check", data=batch_def, suite=suite)
    )
    result = validation.run(batch_parameters={"dataframe": orders})

    for r in result.results:
        cfg = r.expectation_config
        mark = "PASS" if r.success else "FAIL"
        print(f"[{mark}] {cfg.type} {cfg.kwargs.get('column', '')}")
        sample = r.result.get("partial_unexpected_list")
        if not r.success and sample:
            print(f"       unexpected values: {sample}")

    if not result.success:
        print(f"\nContract '{contract['table']}' broken. Owner: {contract['owner']}")
        return 1
    print(f"\nContract '{contract['table']}' holds.")
    return 0


if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

Here's the trick: the exit code is the integration. If the script exits with 1, the CI job fails, and GitHub doesn't need to know anything about GX.

Step 5. The workflow

.github/workflows/data-contract.yml:

name: data-contract

on:
  pull_request:

jobs:
  validate:
    runs-on: ubuntu-latest
    env:
      GX_ANALYTICS_ENABLED: "false"  # no GX usage stats from CI
    steps:
      - uses: actions/checkout@v7
      - uses: actions/setup-python@v7
        with:
          python-version: "3.12"
          cache: pip
      - run: pip install -r requirements.txt
      - name: Validate orders contract
        run: python validate_contract.py
Enter fullscreen mode Exit fullscreen mode

Then, in your repository settings, mark validate as a required status check on main. Otherwise, the check reports the failure but doesn't block the merge.

Let's break it

Run it locally first:

pip install -r requirements.txt
python validate_contract.py
Enter fullscreen mode Exit fullscreen mode

Along with a GX progress bar, you should see every check pass, ending with Contract 'orders' holds.

Now, imagine someone decides .str.lower() looks redundant and removes it:

orders["status"] = orders["status"].str.strip()
Enter fullscreen mode Exit fullscreen mode

Run it again, and among the results you'll see:

[FAIL] expect_column_values_to_be_in_set status
       unexpected values: ['Shipped', 'CANCELLED', 'Delivered']

Contract 'orders' broken. Owner: data-platform@yourcompany.com
Enter fullscreen mode Exit fullscreen mode

The pull request turns red. And look at CANCELLED: since the filter compares against lowercase, cancelled orders are now counted in your totals. That's easy to miss in a one-line diff, yet the contract caught it in about 4 seconds [CONFIRMAR: tempo medido no run do autor].

One last thing

This catches code changes, not surprises from upstream. So keep running the same contract on real data in your orchestrator, and whenever production surprises you, add that row to the fixture.

For more on GX itself, check out this practical guide to automated data validation with Great Expectations on the BIX Tech blog.

Top comments (0)