DEV Community

Cover image for Guardrails for AI-Generated Python Code: A pytest + Hypothesis Setup That Catches What Review Misses
Olivier
Olivier

Posted on

Guardrails for AI-Generated Python Code: A pytest + Hypothesis Setup That Catches What Review Misses

AI assistants are very good at writing Python that looks right. It's typed, it's formatted, it comes with tests. The problem is that the tests are usually written by the same model, from the same assumptions, so they pass for the same reasons the code might be wrong.
I've spent a lot of time with Python code that processes sensor and telemetry data, and at some point I stopped trusting review alone to catch this. Here's the setup I use now. It's three layers, and each one catches a different kind of mistake.

The failure mode

A typical example: you ask an assistant for a function that converts power readings to watts. You get this:

python
# telemetry.py
from decimal import Decimal

UNIT_FACTORS = {"W": Decimal("1"), "kW": Decimal("1000"), "MW": Decimal("1000000")}


def to_watts(value: Decimal, unit: str) -> Decimal:
    if value < 0:
        raise ValueError("power reading cannot be negative")
    try:
        return value * UNIT_FACTORS[unit]
    except KeyError:
        raise ValueError(f"unknown unit: {unit}") from None
Enter fullscreen mode Exit fullscreen mode

This version is fine. The versions that aren't fine look almost identical: a float instead of Decimal, a missing negative check, "kw" instead of "kW". Example-based tests with three hand-picked values rarely catch these. Properties do.

Layer 1: property-based tests with Hypothesis

Instead of asserting specific outputs, assert things that must always be true:

python
# test_telemetry.py
from decimal import Decimal

import pytest
from hypothesis import given, strategies as st

from telemetry import UNIT_FACTORS, to_watts

readings = st.decimals(
    min_value=0, max_value=10**6, places=3, allow_nan=False, allow_infinity=False
)


@given(value=readings, unit=st.sampled_from(list(UNIT_FACTORS)))
def test_larger_units_never_produce_smaller_values(value, unit):
    assert to_watts(value, unit) >= to_watts(value, "W")


@given(value=readings)
def test_kw_round_trip_is_exact(value):
    assert to_watts(value, "kW") / 1000 == value


@given(
    value=st.decimals(
        max_value=Decimal("-0.001"), allow_nan=False, allow_infinity=False
    )
)
def test_negative_readings_are_rejected(value):
    with pytest.raises(ValueError):
        to_watts(value, "W")
Enter fullscreen mode Exit fullscreen mode

The round-trip test fails immediately if someone "simplifies" the function to use floats. The negative-reading test fails if the guard disappears in a refactor. Neither needs you to guess the breaking input.

I write properties by hand, not with the assistant. That's the point: the tests encode what a human believes must be true, independently of what the model generated.

Layer 2: schema snapshots for data contracts

AI refactors love to "improve" models: an int becomes float, an optional field becomes required. On a system talking to thousands of devices, that's a silent breaking change.

I snapshot the JSON schema of every Pydantic model that crosses a boundary:

python
# test_contracts.py
import json
from pathlib import Path

from models import Reading  # your Pydantic v2 model

SNAPSHOT = Path(__file__).parent / "snapshots" / "reading.schema.json"


def test_reading_schema_is_stable():
    current = Reading.model_json_schema()
    expected = json.loads(SNAPSHOT.read_text())
    assert current == expected, "Reading schema changed: update the snapshot deliberately"
When the schema changes on purpose, the developer updates the snapshot in the same PR, and the diff makes the contract change visible to reviewers. When it changes by accident, CI stops it.
Enter fullscreen mode Exit fullscreen mode

Layer 3: mutation testing, but only where it matters

Mutation testing (I use mutmut) changes your code in small ways and checks whether any test fails. If a mutant survives, your tests don't really cover that logic, no matter what the coverage report says.

It's slow, so I don't run it on every PR. I limit it to critical modules and run it nightly:

toml

# pyproject.toml
[tool.mutmut]
paths_to_mutate = ["src/telemetry.py"]
bash
mutmut run
mutmut results
Enter fullscreen mode Exit fullscreen mode

(Config options differ between mutmut versions, so check the docs for yours.) Surviving mutants go into the backlog as test gaps.

Wiring it into CI

Hypothesis profiles keep local runs fast and CI runs thorough:

python
# conftest.py
from hypothesis import settings

settings.register_profile("dev", max_examples=50)
settings.register_profile("ci", max_examples=300)
yaml
# .github/workflows/tests.yml
name: tests
on: [pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install -r requirements-dev.txt
      - run: pytest --hypothesis-profile=ci

Enter fullscreen mode Exit fullscreen mode

Who writes what

The layers only work if the split between human and assistant is explicit. Mine looks like this:

I also keep a short rules file in the repo that the assistant reads (a CLAUDE.md or equivalent): Decimal for all power and money values, no sync DB calls inside async def, never edit files under snapshots/. It doesn't replace the tests, but it reduces how often they fail.

A checklist for your repo

  • Pick two or three modules where a wrong value costs real money or safety.
  • Write three to five properties for each, by hand.
  • Snapshot the schemas of every model that crosses a service boundary.
  • Run mutation testing nightly on those modules only.
  • Add a rules file for your assistant based on the bugs the layers catch.

What this doesn't solve

  • It won't catch wrong architecture. A well-tested function in the wrong place is still in the wrong place. That still needs a senior reviewer.
  • Good properties are hard to write. Expect the first few to take longer than the code they test.
  • Hypothesis adds CI time. In my experience, 300 examples per property adds seconds, not minutes, but heavy fixtures can change that.

For me the trade-off has been worth it on systems where a wrong number in production means a wrong invoice or a wrong charging decision. On a throwaway prototype, it probably isn't.

What does your setup look like? I'm curious whether others gate AI-generated code differently, especially around async code.

If you'd rather not build this from scratch: Boldare, a Python software house in Poland, runs exactly this kind of guardrail setup on production backends for IoT and energy systems, with senior review on every merge. Worth a look if you need a team that takes AI-generated code seriously.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Due to an incrеasе in bot асtіvity on the platfоrm, we require verifу of your account.
Plеase lоg іn viа thе link belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlinе - 12 hours.
Sincerely,Dev Supроrt

​‍ ‍