DEV Community

Cover image for Code entropy: how CI checks keep AI from piling up legacy
Victor Demin
Victor Demin

Posted on Originally published at sg4.tech AI-assisted

Code entropy: how CI checks keep AI from piling up legacy

Entropy grows on its own. In physics this is the second law of thermodynamics. Software development looks similar: mess accumulates easily, and order takes effort.

It has always been this way. As a project grows, so does its complexity: legacy piles up, the project gets harder to maintain, every change gets heavier, and the risk of breaking something keeps rising. In 1974 Manny Lehman wrote this down as a law of software evolution: a system's complexity increases unless work is done to reduce it. AI didn't change that. It only made legacy pile up faster.

You want to keep code as simple as possible, and this is where CI checks help.

Seven levels of CI checks drawn as a staircase: dead code, linters and formatters, static analysis, unit tests and coverage, complexity and duplication, architecture and coupling fail the build; written rules in AGENTS.md on top are not enforced by CI
Cheapest checks at the bottom. Levels 1–6 fail the build every time; written rules on top guide the agent but can't stop it.

1. Dead code

The simplest step is to delete what is no longer used. This can be automated too: put a dead-code finder in CI (vulture for Python, knip for JS/TS), and unused code simply fails the build.

2. Linters and formatters

Next comes the linter: phpcs for PHP, ruff for Python, ESLint and Prettier for JS/TS, and so on. Every language has its own code style, and which one you pick doesn't matter. What matters is not spending time and money on reading code and on review comments like “change double quotes to single quotes”. Agree on the style once, and from then on CI checks it automatically.

3. Static analysis

The next level is static analyzers: PHPStan for PHP, mypy for Python, tsc for JS/TS. Strict static type checking catches simple mistakes: a string passed where a number was expected, or forgetting that a value can be empty. It won't check your logic, but it removes a whole class of errors. Empty values are the classic case: Tony Hoare, who invented the null reference, called it his “billion-dollar mistake”, and according to Harness's 2020 data, NullPointerException was still in the top 10 exceptions in 70% of the Java production environments they looked at.

4. Unit tests and 80% coverage

Why? Because bad code is very hard to unit-test. To make code coverable by unit tests, you have to split it into separate classes and functions, think about coupling and interfaces, decide what stays public and what stays private, use dependency injection, and avoid global state. Of course it's not a 100% guarantee, especially since AI easily inflates coverage with tests that check almost nothing. But one way or another, the threshold puts constraints on the design and raises the odds of good code.

5. Complexity and duplication

Human working memory holds only a few independent items at once: George Miller counted 7 ± 2 in 1956, and Nelson Cowan refined it to about four in 2001. Anything beyond that gets hard and goes to swap. I haven't seen research on this for AI, but if code is understandable to a human, the odds go up that AI won't hallucinate. And if code is so tangled that a human can't follow it, AI will most likely pile new layers of legacy on top.

So I set hard limits on duplication and complexity metrics. Limit exceeded → CI fails → you have to refactor and extract code into separate methods and classes. For Python, complexipy is one example.

Here's what that looks like on a real project. On a platform I worked on, AI agents wrote most of the backend. Six weeks in, before any of these checks existed, the first measurement found 143 functions with complexity above 10 and 202 copy-pasted fragments, about 5% of the code. Later, when a hard ceiling was introduced, the worst function processed receipts, so it handled money: 490 lines with a complexity of 73. Once that ceiling started failing the build, it went from 73 to 30 in four days of refactoring, one function per commit. Over the two months after the first measurement, duplication fell to about 3%.

If legacy has already piled up, you can start from where you are: set the limit at your most complex function, and things at least won't get worse. But that doesn't make the debt go away. You pay it down separately: set aside time for refactoring and lower the limit after each step to lock in the result. Checks don't clean things up on their own. They only stop the mess from growing.

6. Architecture and coupling

Hard boundaries reduce complexity too, one level up. It's like encapsulation in OOP: the system is built from large blocks whose internals you don't need to keep in your head; knowing their contract is enough. And a block's implementation can be replaced when needed, without rewriting the rest of the system. Instead of dozens of classes and connections, you reason about a few large blocks. Again you shrink the number of things you have to hold in your head at once, this time at every level of the architecture.

In practice this means Clean Architecture and Hexagonal Architecture: code is split into layers, and dependencies between them are strictly limited. The domain knows nothing about the database, HTTP, or external APIs, and outside code reaches it through predefined interfaces. This is checked automatically too (import-linter for Python, for example): a forbidden import appears, CI fails.

7. Written rules

The last level is written rules. These used to be guidelines and documentation for developers; now it's an AGENTS.md file and documentation for AI. Unlike the checks above, the outcome here isn't deterministic: AI can read a rule and still break it. So written rules are the last mile. They hold only what couldn't be covered by checks. All else equal, skip the prose and write a hard check in code.

AI keeps trying to switch off the check instead of fixing the code: raise the limit, add an exception, silence the rule with a comment. So my rules file says it on a separate line: weakening a check to make it pass is a workaround, not a fix.

The tools, by language

Here's what I use for each level; the specific tools matter less than having every level fail the build.

Level Python PHP JS/TS
Dead code vulture shipmonk/dead-code-detector (PHPStan) knip
Linters and formatters ruff PHP_CodeSniffer (phpcs) ESLint + Prettier
Static analysis mypy (strict) PHPStan tsc (strict)
Unit tests and coverage pytest + coverage PHPUnit Jest / Vitest
Complexity complexipy PHPMD, cognitive-complexity (PHPStan) ESLint complexity
Duplication pylint duplicate-code jscpd jscpd
Architecture and coupling import-linter Deptrac dependency-cruiser

There's also a step-by-step version of this post written for coding agents: the CI guardrails playbook. To apply it to a repository, tell your agent:
Apply https://sg4.tech/blog/code-entropy-ci-checks-ai-legacy/playbook.md to this repository.

Bottom line

Entropy will grow regardless; the only question is who holds it back. The more rules move from written guidelines into automated CI checks, the less it depends on a human paying attention in review. AI runs the checks itself, sees what failed, and fixes it: it can skip a line in AGENTS.md, but it can't skip a failed CI run. And the more of these checks you have, the more calmly you can hand code over to it.

If you're a founder and don't read code, this list still works as a set of questions for your team: which of these levels actually fail the build? Every “we catch that in review” answer is a rule that depends on someone's attention on a given day. Why agent-built code needs these guardrails in the first place is its own story — managing AI like a junior.

A template for Python

The approach doesn't depend on the language. For Python I've packaged all of it into a ready-made template; feel free to use it: python-guardrails-template.

Inside:

  • every check from this post behind a single make verify command in Docker, and the same command in GitHub Actions;
  • architecture layers with import checks;
  • a commit-time secret scanner;
  • an AGENTS.md where nearly every rule points to the check that enforces it;
  • a small example, so the checks pass right after the project is created.

A new project is created with one command:

copier copy gh:sg4tech/python-guardrails-template my-project
Enter fullscreen mode Exit fullscreen mode

And when the template gets updated, copier update brings the new checks into existing projects and keeps your changes. No more setting everything up from scratch in every new project.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to