A practical look at review capacity, hidden edge cases, and habits that help teams maintain quality.
The principles in this post apply across programming languages. The examples use C# because it is the language I’m most familiar with.
AI-assisted coding has changed the cost of developing software. For many tasks, that is genuinely useful. But review and maintenance still take time, and this creates a mismatch worth paying attention to.
Consider a hypothetical pull request for a payment retry feature. The assistant produces a large, polished change quickly. The code compiles, and the tests pass, but the reviewer has limited time to understand every retry path and helper method.
Weeks later after this feature get's released to production a customer accidently get's charged twice. How could I have missed this?
The risk is not that AI always writes bad code. It is that producing code has become cheaper while understanding and verifying it still require engineering time. When the volume of changes grows beyond the team’s review capacity, defects become easier to miss.
The gap between generation and review
AI coding tools can help developers finish certain tasks faster. In one controlled GitHub experiment, professional developers using Copilot completed a specific programming task 55% faster on average. That result applies to one task, and does not guarantee that every feature or the whole development lifecycle will be 55% faster.1
The risk appears when code production speeds up but review capacity does not.
Illustrative, not measured data
Code arriving: ███████████████
Careful review: ██████
└── The rest waits or gets less attention
When a pull request is too large to understand in the time available, reviewers can end up checking whether it looks right instead of working through what it actually does. Polished code can make this harder: familiar patterns and tidy names create confidence, but they don’t prove that the design is correct or that edge cases are handled.
Faster code generation isn’t the same as faster delivery
DORA’s 2024 research describes a similar tension at the delivery level. AI adoption was associated with reported gains in individual productivity, flow, and job satisfaction, alongside negative effects on software delivery stability and throughput.2
That is not proof that AI inevitably lowers code quality. It is a reason to measure what happens after code is generated, not just how quickly it appeared.
A plausible implementation can hide a subtle edge case
Imagine asking an assistant:
Retry a failed payment request up to three times.
It might generate something like this:
public async Task ChargeOrderAsync(Order order)
{
for (var attempt = 0; attempt < 3; attempt++)
{
try
{
await _paymentGateway.ChargeAsync(order.Amount);
return;
}
catch (Exception) when (attempt < 2)
{
// Retry after a failure.
}
}
}
The loop is easy to follow. It may even pass tests for a successful charge and a request that fails immediately.
But what if the payment gateway processes the charge, then the response times out before the application receives it? The next attempt may charge the customer again.
The problem is not that the code is messy. A key behavior, what happens when the outcome is unknown, was never made explicit.
Make the behavior explicit
A safer design usually needs a stable idempotency key for the logical order, reused across retries. The exact API depends on the payment provider, but the idea might look like this:
public async Task ChargeOrderAsync(Order order)
{
var idempotencyKey = $"order:{order.Id}";
for (var attempt = 0; attempt < 3; attempt++)
{
try
{
await _paymentGateway.ChargeAsync(order.Amount, idempotencyKey);
return;
}
catch (Exception exception)
when (IsRetryable(exception) && attempt < 2)
{
continue;
}
}
}
This is illustrative pseudocode, not a drop-in implementation. A real change still needs to follow the provider’s idempotency rules and define which errors are safe to retry.
Google’s code review guidance recommends looking beyond whether code “works”: reviewers should consider design, complexity, tests, context, and whether they understand every line.3 Those questions matter especially when a patch was generated quickly and looks convincing at a glance.
Practical habits that help protect quality
1. Define the behavior before asking for code
Start with questions, not implementation:
- What should happen when a request times out?
- Which errors should be retried?
- How do we prevent duplicate effects?
- What should happen after the final attempt?
Every AI assistant usually has a plan mode. Ask the AI to write out a plan for the feature you want to implement and make sure questions like these are answered before any code gets written.
2. Keep changes small enough to review
Split work into focused changes: one for the behavior, another for an unrelated cleanup, and another for follow-up documentation.
A small pull request gives a reviewer a better chance to understand how the pieces fit together. It also makes it easier to identify what caused a failure later. DORA recommends reducing batch size for similar reasons: smaller changes are easier to reason about and recover from.4
3. Test the behavior, not just the implementation
For the payment example, useful tests might cover:
- A successful charge.
- A retryable failure.
- A non-retryable failure.
- A timeout after the provider may have processed the charge.
- The same idempotency key being used for every retry.
Don’t stop at “the tests pass.” Ask whether a test would fail if the bug you’re worried about came back. Tests are code too, and a test that repeats the implementation’s assumptions may confirm the wrong behavior.
4. Keep a human responsible for the change
The developer submitting a pull request should be able to explain what every changed part does, why it belongs, and what risks remain.
AI can help summarize a diff or suggest review questions, but that summary should not replace reading the diff. Treat it as a map, not proof that you have visited every place on it.
5. Measure outcomes after generation
Lines generated and tasks started are easy to count; they don’t tell you whether the change helped users or became expensive to maintain.
Look at several signals together: review time, rework, changes that need to be rolled back or urgently fixed, and whether delivery is becoming more or less stable. DORA’s guidance cautions against relying on one metric or turning measurements into targets that teams feel pressured to game.4
These are delivery signals, not direct measures of code quality. Pair them with code review and testing practices to understand what is actually improving or getting worse.
A lightweight review workflow
Define behavior
↓
Ask for a plan
↓
Generate one small change
↓
Test edge cases and inspect the diff
↓
Human review and ownership
↓
Merge, monitor, and learn
Before approving, ask:
- Do I understand every meaningful change?
- What happens on a timeout, duplicate request, or partial failure?
- Would the tests catch the failure we care about?
- Is this change smaller and simpler than it could be?
- Can the author explain the tradeoffs?
If the honest answer to “Do I understand this?” is no, ask for clarification or a smaller change before approving it.
Keep understanding in the loop
AI can shorten the distance between an idea and a working first draft. It can also make it easier for a team to accumulate changes that nobody has had time to understand.
The useful goal is to keep the size and arrival rate of changes within the team’s ability to review them carefully. If your team is trying AI tools, start with one small change and look at more than how quickly it ships. Check the tests, review effort, and cost of maintaining the result too.
Sources and further reading
- GitHub: Research quantifying GitHub Copilot’s impact on developer productivity and happiness: includes the 55% result from a controlled experiment on a specific programming task.
- DORA: 2024 Accelerate State of DevOps Report: reports benefits and tradeoffs associated with AI adoption.
- Google Engineering Practices: What to look for in a code review: guidance on reviewing design, functionality, complexity, tests, and context.
- DORA: Software delivery performance metrics: explains delivery throughput and stability measures, including why small batches help.
- METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity and METR’s February 2026 study update: useful context for why AI’s productivity impact depends on the tools, tasks, and study design. METR’s update cautions that selection effects made its newer estimate unreliable.
-
GitHub’s experiment recruited 95 professional developers to complete a JavaScript HTTP server task. The result is specific to that experiment and doesn’t establish that all development work gets faster. ↩
-
DORA reports associations from its research; those findings don’t establish that AI alone caused changes in delivery performance. ↩
-
Google’s guidance was written for code review generally, not specifically for AI-generated changes. ↩
-
DORA presents these as delivery performance signals, not direct measures of code quality. ↩
Top comments (0)