A reviewer finds a bug. A coding agent fixes it. The fix triggers another review. The reviewer finds a style issue. The agent changes it. A new run notices a nearby concern. Soon the pull request has ten commits, three of them correcting earlier corrections, and nobody is sure which result is the one to review.
That loop is not proof that AI review is useless. It is a sign that the team has not defined a stopping rule.
Review should converge on a decision
Automated review can produce many observations, but a pull request still needs a finite decision: accept the change, request a bounded correction, or stop because evidence is missing.
A practical pattern is one correction wave:
- Collect. Let the intended reviewers and required checks finish for the current revision.
- Normalize. Group duplicate comments and distinguish actionable defects from questions, style preferences, and speculation.
- Triage. Assign an owner and consequence level. Validate the failure path before changing code.
- Correct coherently. Make a related set of accepted fixes together instead of committing each bot suggestion separately.
- Re-run evidence. Run required deterministic checks on the new head and inspect the resulting diff.
- Decide. Merge, request a specifically scoped next wave, or stop with the remaining unknowns named.
This is a process pattern, not an experimentally proven optimum. The useful property is that it makes the loop observable and gives it a natural point to end.
Why serial micro-fixes create noise
Every new commit changes the object under review. A reviewer may now be commenting on a revision that no longer exists. A small patch can change the meaning of a previous test. A bot may rerun because the head SHA changed and repeat a finding that was already dismissed.
Serial fixes also fragment intent. A single change to validation, error handling, and a regression test may be easier to evaluate as one cohesive unit than as a chain of tiny commits whose dependency is implicit.
The opposite extreme is not better. A large “fix everything” commit can hide unrelated edits. Coherence means the corrections share one accepted invariant or failure path, not that every review comment is bundled indiscriminately.
A triage card for each finding
For every material finding, capture:
- Claim: what can go wrong?
- Path: how does the code reach the failure?
- Evidence: what test, reproduction, or source confirms it?
- Owner: who decides whether it is in scope?
- Disposition: accept, reject with reason, or defer with an owner.
- Durable lesson: should the fix become a regression test or policy check?
This card can be brief. Its purpose is to stop a confident comment from turning directly into a code mutation.
When a second wave is justified
A new correction cycle makes sense when it starts with new evidence: a required test fails on the corrected revision, a reviewer identifies a distinct material defect, or a reproduction changes the risk assessment.
It does not make sense merely because a bot can be triggered again. Re-run a check when the input changed or when the check did not cover the intended revision. Do not ask the same nondeterministic reviewer to rediscover a finding until it returns a satisfying answer.
Define the next wave’s scope before opening it:
- what new evidence triggered it;
- which findings are in scope;
- what artifact or revision is under review;
- which checks must pass;
- what condition ends the loop.
If a review provider is flaky or unavailable, represent that state honestly. “No comment” may mean no issue was found, or it may mean the review never completed. The status should distinguish those cases.
Keep liveness in the design
Safety asks whether a bad change can get through. Liveness asks whether a good change can get stuck forever.
An AI system that can keep creating commits, re-triggering itself, or reopening dismissed findings can threaten both. Limit automatic retries. Keep comments tied to a precise revision. Avoid granting the reviewer credentials that let it silently approve its own fixes. Preserve a human-visible way to terminate the cycle.
None of these controls guarantee a fast review. They make it possible to see why the review is waiting and who can move it forward.
Measure the loop without gaming it
Count more than comments or generated commits. A team might track:
- time from first review to a merge decision;
- distinct material findings validated;
- accepted findings that led to a regression guard;
- repeated findings on unchanged behavior;
- corrections that had to be reverted;
- checks skipped, stale, or inconclusive.
These measurements need context. A lower comment count is not automatically better; a shorter review can mean a missed defect. Use a portfolio view and pair speed with quality and rework.
A review protocol should terminate
Good review is not maximum review. It is enough review to support a decision, with unresolved risks made visible.
Batching accepted findings into one coherent correction wave helps keep the diff legible and the state machine finite. Re-run deterministic evidence against the corrected revision. If meaningful new evidence appears, open another bounded wave. If the evidence has not changed, stop repeating the same prompt.
The next post follows one accepted finding beyond the pull request: how it becomes a durable deterministic guard, instead of a note everyone forgets after merge.
References
- GitHub: About status checks
- GitHub Actions security reference
- DORA, Balancing AI tensions (2026) — qualitative analysis of engineer responses, not a controlled experiment establishing this exact workflow.
AI assistance was used to prepare this draft. The human editor is responsible for validating the workflow advice against the intended repository and reviewing the final text.

Top comments (0)