DEV Community

Cover image for How AI Can Lead to a Decline in Code Quality and How to Fix It
Mitar
Mitar

Posted on

How AI Can Lead to a Decline in Code Quality and How to Fix It

A practical look at review capacity, hidden edge cases, and habits that help teams maintain quality.

The principles in this post apply across programming languages. The examples use C# because it is the language I’m most familiar with.

AI-assisted coding has changed the cost of developing software. For many tasks, that is genuinely useful. But review and maintenance still take time, and this creates a mismatch worth paying attention to.

Consider a hypothetical pull request for a payment retry feature. The assistant produces a large, polished change quickly. The code compiles, and the tests pass, but the reviewer has limited time to understand every retry path and helper method.

Weeks later after this feature get's released to production a customer accidently get's charged twice. How could I have missed this?

The risk is not that AI always writes bad code. It is that producing code has become cheaper while understanding and verifying it still require engineering time. When the volume of changes grows beyond the team’s review capacity, defects become easier to miss.

The gap between generation and review

AI coding tools can help developers finish certain tasks faster. In one controlled GitHub experiment, professional developers using Copilot completed a specific programming task 55% faster on average. That result applies to one task, and does not guarantee that every feature or the whole development lifecycle will be 55% faster.1

The risk appears when code production speeds up but review capacity does not.

Illustrative, not measured data

Code arriving:  ███████████████
Careful review: ██████
                └── The rest waits or gets less attention
Enter fullscreen mode Exit fullscreen mode

When a pull request is too large to understand in the time available, reviewers can end up checking whether it looks right instead of working through what it actually does. Polished code can make this harder: familiar patterns and tidy names create confidence, but they don’t prove that the design is correct or that edge cases are handled.

Faster code generation isn’t the same as faster delivery

DORA’s 2024 research describes a similar tension at the delivery level. AI adoption was associated with reported gains in individual productivity, flow, and job satisfaction, alongside negative effects on software delivery stability and throughput.2

That is not proof that AI inevitably lowers code quality. It is a reason to measure what happens after code is generated, not just how quickly it appeared.

A plausible implementation can hide a subtle edge case

Imagine asking an assistant:

Retry a failed payment request up to three times.

It might generate something like this:

public async Task ChargeOrderAsync(Order order)
{
    for (var attempt = 0; attempt < 3; attempt++)
    {
        try
        {
            await _paymentGateway.ChargeAsync(order.Amount);
            return;
        }
        catch (Exception) when (attempt < 2)
        {
            // Retry after a failure.
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

The loop is easy to follow. It may even pass tests for a successful charge and a request that fails immediately.

But what if the payment gateway processes the charge, then the response times out before the application receives it? The next attempt may charge the customer again.

The problem is not that the code is messy. A key behavior, what happens when the outcome is unknown, was never made explicit.

Make the behavior explicit

A safer design usually needs a stable idempotency key for the logical order, reused across retries. The exact API depends on the payment provider, but the idea might look like this:

public async Task ChargeOrderAsync(Order order)
{
    var idempotencyKey = $"order:{order.Id}";

    for (var attempt = 0; attempt < 3; attempt++)
    {
        try
        {
            await _paymentGateway.ChargeAsync(order.Amount, idempotencyKey);
            return;
        }
        catch (Exception exception)
            when (IsRetryable(exception) && attempt < 2)
        {
            continue;
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

This is illustrative pseudocode, not a drop-in implementation. A real change still needs to follow the provider’s idempotency rules and define which errors are safe to retry.

Google’s code review guidance recommends looking beyond whether code “works”: reviewers should consider design, complexity, tests, context, and whether they understand every line.3 Those questions matter especially when a patch was generated quickly and looks convincing at a glance.

Practical habits that help protect quality

1. Define the behavior before asking for code

Start with questions, not implementation:

  • What should happen when a request times out?
  • Which errors should be retried?
  • How do we prevent duplicate effects?
  • What should happen after the final attempt?

Every AI assistant usually has a plan mode. Ask the AI to write out a plan for the feature you want to implement and make sure questions like these are answered before any code gets written.

2. Keep changes small enough to review

Split work into focused changes: one for the behavior, another for an unrelated cleanup, and another for follow-up documentation.

A small pull request gives a reviewer a better chance to understand how the pieces fit together. It also makes it easier to identify what caused a failure later. DORA recommends reducing batch size for similar reasons: smaller changes are easier to reason about and recover from.4

3. Test the behavior, not just the implementation

For the payment example, useful tests might cover:

  • A successful charge.
  • A retryable failure.
  • A non-retryable failure.
  • A timeout after the provider may have processed the charge.
  • The same idempotency key being used for every retry.

Don’t stop at “the tests pass.” Ask whether a test would fail if the bug you’re worried about came back. Tests are code too, and a test that repeats the implementation’s assumptions may confirm the wrong behavior.

4. Keep a human responsible for the change

The developer submitting a pull request should be able to explain what every changed part does, why it belongs, and what risks remain.

AI can help summarize a diff or suggest review questions, but that summary should not replace reading the diff. Treat it as a map, not proof that you have visited every place on it.

5. Measure outcomes after generation

Lines generated and tasks started are easy to count; they don’t tell you whether the change helped users or became expensive to maintain.

Look at several signals together: review time, rework, changes that need to be rolled back or urgently fixed, and whether delivery is becoming more or less stable. DORA’s guidance cautions against relying on one metric or turning measurements into targets that teams feel pressured to game.4

These are delivery signals, not direct measures of code quality. Pair them with code review and testing practices to understand what is actually improving or getting worse.

A lightweight review workflow

Define behavior
      ↓
Ask for a plan
      ↓
Generate one small change
      ↓
Test edge cases and inspect the diff
      ↓
Human review and ownership
      ↓
Merge, monitor, and learn
Enter fullscreen mode Exit fullscreen mode

Before approving, ask:

  • Do I understand every meaningful change?
  • What happens on a timeout, duplicate request, or partial failure?
  • Would the tests catch the failure we care about?
  • Is this change smaller and simpler than it could be?
  • Can the author explain the tradeoffs?

If the honest answer to “Do I understand this?” is no, ask for clarification or a smaller change before approving it.

Keep understanding in the loop

AI can shorten the distance between an idea and a working first draft. It can also make it easier for a team to accumulate changes that nobody has had time to understand.

The useful goal is to keep the size and arrival rate of changes within the team’s ability to review them carefully. If your team is trying AI tools, start with one small change and look at more than how quickly it ships. Check the tests, review effort, and cost of maintaining the result too.

Sources and further reading


  1. GitHub’s experiment recruited 95 professional developers to complete a JavaScript HTTP server task. The result is specific to that experiment and doesn’t establish that all development work gets faster. ↩

  2. DORA reports associations from its research; those findings don’t establish that AI alone caused changes in delivery performance. ↩

  3. Google’s guidance was written for code review generally, not specifically for AI-generated changes. ↩

  4. DORA presents these as delivery performance signals, not direct measures of code quality. ↩

Top comments (0)