Why Code Review Suggestions Never Get Implemented
Last month I caught myself saying the same sentence for the third time in a code review:
"Let's not fix this now — we'll keep it in mind next time."
First time was three months ago. Second time was last month. Third time was now.
Every time: "next time." Every time: nobody remembered. Including me.
So I went back and looked at the review comments from the past six months. The result was sobering: most suggestions were never implemented.
Not rejected. Not debated. Just... forgotten.
"Next time" is a polite way of saying "never"
Let's take that sentence apart:
"Let's not fix this now — we'll keep it in mind next time."
It sounds courteous. It keeps the meeting civil. But it does three things at once:
- It gives the other person permission not to act.
- It gives me a reason not to insist.
- It erases the suggestion entirely — no record, no owner, no deadline.
So the practical meaning is: not now, not next time, not ever.
I call this politely useless. Everyone feels fine. Nothing gets fixed.
Why verbal reminders can't work
Think about what has to happen for a review suggestion to actually get implemented:
| Step | What actually happens | Result |
|---|---|---|
| Suggestion is raised | Said out loud in a meeting | No artifact |
| It's recorded | Nowhere, or in someone's head | Usually lost |
| It's assigned | No owner, no deadline | Nobody's job |
| It's tracked | Nobody looks again | Dies quietly |
| It's implemented | Depends on good intentions | Depends on the day |
Every one of these steps is broken.
If any single link fails, the suggestion is dead. And with verbal reminders, no link is connected at all.
The deeper problem: we're using human memory as a task queue. Human memory has never been good at that. A sentence like "next time" survives maybe ten minutes in your head.
The most common excuse: "No time, the sprint is tight"
The analysis above is a bit idealized. The version you actually hear is:
"I know it needs fixing, but we're underwater on this sprint. Let's revisit it later."
I'll grant that this is partly true. There are deadlines, quarters, OKRs. A "style" problem that doesn't affect functionality will always sort to the bottom.
But here's the thing: "later" never arrives.
The typical lifecycle of such a suggestion:
| Time | State |
|---|---|
| Review day | "Noted, I'll fix it later" |
| One week later | Shipped something, firefighting, forgotten |
| One month later | Fully forgotten; code is in production, so fixing is now riskier |
| Three months later | The same person writes the same mistake again |
That last row is the real cost.
"No time to fix it" doesn't make the problem go away. It makes it reappear later, in more places, at higher cost.
The half hour you saved today will come back at ten times the price — and not just once.
So "the sprint is tight" is not a reason to skip it. It's the reason to let a machine handle it:
- Rely on people → always behind the sprint, always "no time"
- Rely on the build → it blocks at commit time, takes zero human time
If a machine can fix it automatically, don't spend a human's sprint time on it.
A deeper problem: we review at the wrong time
Everything above is about what happens after a suggestion is made. There's a more fundamental issue.
A common workflow is: write the code, then review it — and review the technical design in the same pass.
This sounds thorough. It's actually a process bug:
The design should have been settled before anyone wrote code.
In practice:
| Phase | What should happen | What actually happens |
|---|---|---|
| Requirements → Design | Design review; architecture and key decisions settled | Skipped, or a verbal "sounds good" |
| Design → Coding | Implement against the design | Design lives in one person's head |
| Coding → Review | Verify implementation matches the design | Design and implementation reviewed together |
See the problem?
If you're discussing the design while reviewing the code, the design was decided in the code — and the code was already written.
Two consequences:
1. Design mistakes get held hostage by sunk cost.
Changing a design before coding means editing a diagram. Changing it after coding means rewriting. So people compromise: "it's already written, let's leave it." The mistake gets baked in permanently.
2. One review, two goals, both done badly.
Design review needs divergence — debate, trade-offs, exploration. Implementation review needs convergence — check against a standard. Mixing them in one meeting means neither is done properly — and whatever got missed becomes the next round of "next time."
The review is in the wrong place. It shouldn't be the safety net for design and implementation. It should answer one question: does the implementation match the agreed design?
The fix: review in stages
Review shouldn't be a single event — it should be staged according to the size of the change:
| Stage | What to review | When |
|---|---|---|
| 1 | Overall structure | Once the skeleton exists, before filling in details |
| 2 | Core logic | Once the main implementation is done |
| 3 | Details and conventions | Once it's functionally complete |
Why this works:
- Problems get caught when they're cheapest to fix. A structural problem fixed at stage 1 is a few class moves. Found at the end, it's a rewrite.
- Each review has a single goal. Stage 1 debates structure, not naming. Stage 3 polishes details, not architecture. Reviewers stay focused.
- Large changes become executable. Sometimes "next time" isn't laziness — the change is just too big to do in one pass. Staged review keeps each increment small.
Review cadence is itself a convention. The bigger the change, the more stages it needs. Saving it all for one big review guarantees a backlog of "next time."
The question isn't "do we fix this now?" It's "at which stage should this have been caught?"
The root cause: we manage machines' work with human reminders
Here's what I think the real problem is:
We use human processes for things that should be mechanical.
Review suggestions fall into two categories.
Category 1: machine-detectable (the majority)
Naming, method length, nesting depth, logging, layering violations (a controller calling a DAO directly), duplicated code, inconsistent exception handling.
All of this is automatable. Yet we say it out loud, expect someone to remember it, and hope they'll fix it voluntarily.
That's an enormous waste. A machine can do this ten thousand times a second. We do it with a meeting, three sentences, and one person's memory.
Category 2: not machine-detectable (the minority)
Whether a name reflects the domain, whether an abstraction earns its weight, whether logic belongs in that layer.
These genuinely need a human. But even these shouldn't be handled with "next time" — they should become trackable items.
What actually works
Rule 1: If a machine can block it, don't say it out loud.
Write every "we must remember this from now on" rule as a check, wire it into CI, and block the merge if it fails.
- Formatting → auto-format on commit, fail the build if dirty
- Static analysis → quality gate, block merge if new issues exceed the threshold
- Architecture → encode rules like "no cross-layer calls" as executable tests
The key is not detection. It's blocking. A check that only warns is equivalent to no check, because someone will click "ignore."
Warning without blocking is politely useless.
Rule 2: If a machine can't judge it, make it a trackable item.
A "won't fix now" suggestion must immediately become a ticket with an owner and a due date. Otherwise it's just a compliment.
- Every suggestion goes into the review system's comments — no more verbal notes
- Must-fix items become blocking, and the merge is gated on them
- Won't-fix-now items create a linked ticket, assigned, with a date
Three months later, every suggestion has a traceable history.
Rule 3: Repeated suggestions should be promoted to rules.
This is the most important one.
If you've made the same suggestion three times, the problem isn't the person — it's your process.
A suggestion raised three times should be promoted from "verbal advice" to "automated check." Stop reminding people. Let the machine say it for the hundredth time.
The quieter loss: handover wipes it all out
Everything above assumes the person who raised the suggestion is the person who fixes it.
But people leave. Teams reorganize. Modules change hands. When that happens, every problem above gets worse.
A verbal suggestion's fate during handover:
| Scenario | What happens to the suggestion |
|---|---|
| Same person, same team | At least a chance of "next time" |
| Code handed to someone else | Gone |
| Handover + original author leaves | The rationale is gone too |
First: suggestions don't transfer.
"Let's remember this next time" exists only inside that one meeting. It was never written down, assigned, or tracked. When the code changes hands, it isn't in any handover document.
The new owner doesn't even know the suggestion existed. That's not carelessness — the information was never captured.
Second: tech debt gets silently inherited.
Worse than losing a suggestion: the previous owner's debt gets inherited under the label "legacy." The new owner looks at confusing code and thinks:
"There's probably a historical reason. I don't understand why it's written this way, so I'll leave it alone."
Then adds features on top of it.
The debt isn't repaid — it becomes the foundation for new code. The longer it sits, the less anyone dares touch it.
Third: the most expensive loss — rationale disappears.
The code is visible. "Why it's designed this way" lives only in the original author's head:
- A cleaner approach was avoided because of a specific failure mode
- A seemingly redundant check was added after an incident
- A "needs refactoring" spot hides an unresolved problem
None of this is in the code. None of it is in the handover doc.
Two failure modes follow: someone tears out a design that had good reasons and rediscovers the original bug — or preserves a broken implementation believing it's best practice.
Code can be handed over. Intent, if it isn't captured, leaves with the person.
AI closes the loop — but not how most teams think
This methodology has always had one blocker: cost.
- Writing checks is expensive, especially architectural constraints
- Turning suggestions into tickets is manual, and manual means forgotten
- Judging which suggestions deserve to become rules requires experience
These costs are why "make it a rule" stayed a slogan. Writing rules is tedious. Saying it out loud is one sentence.
Recent progress in coding LLMs and agents has collapsed most of these costs.
1. Rule generation is now nearly free.
Writing an architecture rule used to mean digging through APIs and debugging a test. Now you can say:
"Scan this module for controllers that depend directly on DAOs, and generate a CI-ready architecture rule."
The model understands your codebase and produces a runnable rule. The "three strikes and it becomes a rule" policy is finally practical.
2. Agents make staged review automatic.
Staged review needs someone at every stage — and reviewers are busy. Agents fill the gap: structure and dependency direction at stage 1, boundaries and error handling at stage 2, naming and formatting at stage 3.
Most importantly: they separate "newly introduced" from "pre-existing." The biggest problem with human review is that legacy noise buries new issues. An agent can look only at the diff and flag what this change introduced — exactly what "next time" tends to miss.
3. Suggestions become actionable.
"This method is too long" is abstract. A model can tell you which methods to extract, what to call them, which call sites are affected, and what priority to assign.
Suggestions stop being reminders and become tasks. Part of "no time to fix it" was that suggestions were too vague and too expensive to start.
4. The best part: AI itself can be bound by your conventions.
Feed your team's conventions to the model and it will follow them while generating code. Instead of post-hoc checks, you get alignment at generation time.
From blocking after the fact to aligning before the fact.
But AI does not replace process
Many teams assume that adopting an AI code assistant automatically improves code quality.
It doesn't.
- AI raises a suggestion, and nobody tracks it → still ignored
- AI finds a problem, and it doesn't block the merge → someone clicks "ignore"
- AI raises it every single time, and nobody promotes it to a rule → it'll raise it again tomorrow
AI reduces the cost of finding and solving. It doesn't reduce the cost of executing.
AI handles detect + generate + explain. Process handles block + track + measure.
You need both:
- People without AI → slow, incomplete, relies on goodwill
- AI without process → more suggestions, none executed, more noise
- AI + process → rules generated, merges blocked, debt tracked; humans make the final call
A counterintuitive observation: in the AI era, code review becomes more valuable, not less. Generation speed has gone up dramatically. The judgment of whether code should be written that way still requires a human. The bottleneck moved from writing to reviewing.
The counterintuitive conclusion
The value of a code review isn't finding new problems. It's finding conventions that should be automated.
A review that finds a problem, mentions it, and ends has near-zero value — the same problem will recur.
But a review that produces either a new automated check or a trackable ticket raises the team's floor permanently.
The goal isn't to get this fixed once. It's to never have to say it again.
Where I'm still stuck
Honestly, I haven't fully solved this. The methodology above is still expensive to run:
- Writing checks costs effort; architectural constraints especially
- Turning every "next time" into a ticket requires tooling, or humans forget
- Rules end up scattered across CI configs, linters, and issue trackers — nowhere to see "which conventions have we actually fought about"
My conclusion: there's a missing piece between "rules a machine can enforce" and "suggestions a human must track" — something lightweight that connects them.
I doubt I'm the only one.
How does your team handle this?
If you're a tech lead or you run code reviews, I'd genuinely like to know:
- Roughly what fraction of your review suggestions actually get implemented?
- How do you handle "won't fix now, but remember next time"?
- Is there a convention you've raised a hundred times and people still break?
Curious whether this is a universal problem or just my experience.
Top comments (1)
We've had cases on our team where code-review suggestions were ignored for long while, resulting in a buildup of technical debt.....