An automatic code reviewer now runs on my blog repository — OpenAI's Codex review, which
kicks in on its own when a PR opens. The first review in the record is dated 7 October. It
writes what it finds line by line, and reacts with 👍 when it finds nothing. Not a bad
deal: a second pair of eyes on a one-person repo, and free at that.
Today I got curious: what did this reviewer say, and what did I do about it? I went
through the records. In four of five pull requests the reviewer had something to say. In
three of them I had hit the merge button before it finished reading. The findings landed
against a closed door: one written thirteen seconds after the merge, one seventy-seven
seconds after, two a hundred and one seconds after.
I did catch one of the four anyway — in four and a half minutes, fixing it inside the next
pull request. The other three are still sitting on the live site today.
This piece is the write-up of that measurement. The interesting part isn't "I ignored the
reviewer," because I didn't. The interesting part is that which findings I heard was
left entirely to chance: to whether I still happened to be looking at that PR. Because
what makes a gate a gate is not what it produces but what it can hold back — and mine
didn't hold back a single merge.
What good is a pull request that merges in seven seconds
Start with the wide view. The blog repository moved to its current home on 14 September
- Since then 228 commits have landed on the first-parent line of the main branch. Seventeen of them came through a pull request. The remaining 211 were written straight to main — no PR, no reviewer, no gate to wait at.
So my PR flow covers roughly one thirteenth of the work that reaches main. The rest is
handled by git push origin main. I already knew that much; the content pipeline writes
articles directly to main, and I branch for design and infrastructure work. The
uncomfortable part is how long those seventeen pull requests lived.
Here are the open-to-merge times, in seconds:
7 8 9 11 17 19 21 21 24 27 103 141 157 187 334 726 741
Median: 24 seconds. Ten of the seventeen closed in under a minute. The quickest was
seven seconds. The longest was a sidebar PR with "rev3" in its title: 12.3 minutes.
What happens in seven seconds? The page refreshes, the button goes live, you press it.
That is not a review, it's a bookkeeping style: I branched the work, I wrote the branch
into main, let the record show it. The value isn't zero — reverts get easier, commits get
grouped. But there's nothing in it that justifies the phrase "I reviewed this."
Until early October none of this mattered, because the first twelve pull requests had
neither checks nor a reviewer. Across those twelve PRs, from 14 to 23 September, the
check count is zero and the review count is zero. Walking through an empty doorway in
seven seconds is no worse than walking through it slowly.
The real question is what I did once I put an actual guard at the door.
There was no door at all
Before I describe the guard, I should say there was never a wall. I asked the API about
the protection settings on main. The answer is short:
GET /repos/.../branches/main/protection
→ 404 {"message": "Branch not protected"}
In the repository settings, allow_auto_merge is false as well. So two things are true
at once: nothing on main requires any check or approval, and the mechanism that would
queue a merge behind the checks is switched off.
GitHub's own documentation ties those two together. For an unprotected branch: "If
required status checks aren't enabled, collaborators can merge the branch at any time,
regardless of whether it is up to date with the base branch." And for auto-merge:
"Auto-merge merges a pull request automatically after all required reviews and status
checks pass" — an option that only appears when a pull request cannot merge immediately,
which is to say when some protection rule is holding it back.
There's a neat circularity here. Auto-merge is the tool that automates waiting. But if
there is nothing to wait for, auto-merge never even shows up on screen. With no
protection, "waiting" isn't an institution — there is only whether I happen to be in a
hurry.
I can compress this article's thesis into one line: a gate nothing enforces is not a
gate but a habit, and habits lose to haste.
7 October: the reviewer read five times, I failed to wait three times
On 7 October I rewrote the interface and opened five pull requests. By then both the
CodeQL checks and the reviewer were active. On each PR the reviewer leaves a "Review
Summary" comment when the review starts and updates that same comment to "Completed" when
it ends. So the record holds both a start and an end time. GitHub keeps the merge time.
Lining up all three was enough.
| PR | Review started | I merged | Review finished | Findings |
|---|---|---|---|---|
| #21 | 14:35:38 | 14:47:34 | 14:40:33 | 1 × P1 |
| #22 | 15:13:12 | 15:15:35 | 15:16:54 | 1 × P2 |
| #23 | 15:25:54 | 15:31:15 | 15:28:24 | none (👍) |
| #24 | 15:32:07 | 15:35:01 | 15:35:17 | 1 × P2 |
| #25 | 15:51:50 | 15:54:00 | 15:55:43 | 2 × P2 |
I stared at this for a while the first time I laid it out. On #22, #24 and #25 the merge
time comes before the review's finish time. I shut the door while the reviewer was still
reading. And this wasn't invisible: the review's record was sitting right there on the
PR page at the time, and the "completed" stamp landed after I merged.
The gaps are small. Thirteen seconds on #24. I could have fetched a coffee and still made
it.
Where those four findings are today
The three late reviews produced four findings between them. All P2 — the "not broken but
wrong" class. I checked each of them today, against the source and against the live
site.
1. Accumulating listeners (#22) — fixed, but not thanks to the gate. The reviewer
pointed out that on every client-side navigation the top bar's menu code attached two
fresh document-level listeners and never removed the old ones. I did read this one: it
arrived at 15:16:52 and I pushed the fix commit at 15:21:26 — four minutes
thirty-four seconds. The commit message names the defect outright: "menu close listeners
are attached to document ONCE, finding the menu at event time; previously every
ClientRouter navigation added 2 new listeners." Underneath it there's even a verification
note showing the listener count held steady across four navigations.
So how long did the defect live on main? Because I put the fix inside a separate pull
request (#23) and merged that at 15:31:15: 15 minutes 40 seconds. The system worked
here. But note what worked — it wasn't the gate. I happened to still be looking at the
same file and I saw the notification. The gate couldn't have stopped anything; it was
already open.
2. Turkish status strings on English pages (#24). The personal-notes box under each
article picks its opening status by page language; but once you start typing it says
"Kaydediliyor...", then "Kaydedildi" or "Kaydedilemedi" — regardless of language. The
reviewer wrote this thirteen seconds after my merge. It is still true today. I opened a
live English article and looked: lang="en", opening status "Ready", and the three
Turkish strings sitting intact inside the script shipped to the page.
3. The entrance animation overriding hover (#25). This became my favourite finding,
because it caught the defect inside the very PR that introduced it, before the ink was
dry. I gave the panels a "rise into view" animation and set its fill mode to both.
Right underneath I wrote .tm-pane:hover { transform: translateY(-2px) }, so a panel
would lift slightly under the cursor. The reviewer said that once the animation ends, the
final keyframe's translateY(0) stays applied to the element, and animation-sourced
values override normal rules — so the hover lift never runs at all.
I checked this on the live site. I hovered a panel and read the computed value:
:hover matches? true
animation fx-rise : finished : fill=both
computed transform matrix(1, 0, 0, 1, 0, 0)
expected (hover) matrix(1, 0, 0, 1, 0, -2)
The cursor is on the panel, the browser is matching the :hover rule, and nothing
happens. Two pixels. Two pixels nobody will complain about and I would never have gone
looking for.
The reviewer was right, and the specification says why. CSS Animations is explicit:
"animations override all normal rules, but are overridden by !important rules." MDN's
definition of the forwards fill mode (which both includes) is just as clear: "The
target will retain the computed values set by the last keyframe encountered during
execution." Put those side by side and the outcome is unavoidable: a retained animation
value beats a normal hover rule every time. There's no browser quirk here; I simply
hadn't read the rule.
Poking at this finding turned up another one beneath it: further down the same file sits
.tm-pane { transition: border-color 180ms ease; }, and that shorthand overrides the
transition: transform 0.35s declaration above. So even with the animation problem
solved, the panel would jump rather than glide. A finding under a finding.
4. Dark-theme code blocks in print (#25). I wrote rules that force code-block colours
in dark theme with !important. The print stylesheet touches only the border of a code
block, not its colour or background. Since !important wins in print too, anyone
printing an article with dark theme active gets either dark-theme colours on white paper
or an ink-devouring block. This one is still true as well.
To summarise: four findings, one fixed in four and a half minutes, three live. Not one of
them has a 👍 underneath it, or a reply. The only trace of my having heard one is a commit
message.
What separated the ones I heard from the ones I didn't
This is the part that stings, because I have my own control group — and what did not
separate them is very clear.
It wasn't severity. Of the two findings I acted on, one was a P1 and one a P2. The
three I missed were all P2. Had the reviewer written a P1 on any of those three PRs, the
same thing would have happened — the gate doesn't distinguish severity, because there is
no gate to do the distinguishing.
It wasn't diligence either. On #21 the reviewer found a P1: I had written an
unverified, site-wide measurement claim into the new homepage copy. The review finished at
14:40:33; seventy seconds later I pushed the fix commit, and at 14:41:50 I replied under
the comment saying "fair catch." On #22 I fixed it in four and a half minutes. When I'm
told, I fix it; that isn't the problem.
The only thing that separated them was whether I was still looking at that PR when the
notification dropped. On #21 twelve minutes passed between opening and merging, and the
finding landed in the middle of that window. On #22 I had merged but hadn't yet left the
page. On #24 and #25 I had already moved to the next job; one finding was thirteen seconds
late, the other a hundred and one, and I wasn't there.
So the process works. The reviewer really does read, really does find things, and I really
do fix them when I'm told. The only missing piece is the merge waiting — and that missing
piece decides, by coin toss, which findings get to live.
CI passed, but that is not evidence
I could have consoled myself with "the CodeQL checks always passed." They did. Across all
five PRs every check succeeded, and every one of them completed before the merge.
Sounds like a working gate.
Until you look at the gap between the last check finishing and my merge:
| PR | Last check finished → merge |
|---|---|
| #25 | 6 seconds |
| #22 | 28 seconds |
| #24 | 40 seconds |
| #21 | 163 seconds |
| #23 | 201 seconds |
Six seconds. I have a "passed its checks" commit only because CodeQL happened to be six
seconds faster than me that time. Had it had a slow day, or had I read one paragraph
less, that commit would have gone in unchecked — and nothing would have stopped it.
The second problem is simpler: across those five PRs no check ever failed. The one
experiment that would show whether the door can close has never been run. I don't have
evidence that the gate works; I have an observation that the gate has never needed to
work. Those aren't the same thing. Sixty-eight configuration copies under
/etc made the
same point six days ago: a copy that exists and a copy that does anything are separate
line items.
My own defence, and where it collapses
Let me be fair to myself, because it's possible to look at these findings and shrug. All
four are P2. Nobody hit a broken page, no data was lost, the site works. On a one-person
blog, fixing a two-pixel hover effect a day late is not the world's most expensive
mistake. Speed is a real value; I rewrote the interface that day and only finished because
I could fit five PRs into an hour and a half.
The defence holds up to a point. Here's where it collapses: in those three merges I never
decided "this is a P2, I'll deal with it later." I made no decision at all. I pressed
the button without seeing the finding. The outcome — three small defects going live —
looks like the result of a reasonable trade-off, but it isn't; it's a residue of timing.
That's the difference between taking a risk knowingly and arriving at the same place
without noticing. The first is an engineering decision, the second is a coin toss. The
outputs sometimes match — which is exactly why telling them apart requires measurement.
Why "I'll be more careful" isn't a fix
My first instinct writing this was to promise myself something: from now on I won't merge
before the reviewer finishes. That'll hold for a day. Maybe a week. Then it'll be eleven
at night, a small CSS correction, the checks will look green, and six seconds will be six
seconds again.
The genuinely embarrassing part is that I already knew all this. In the piece about the
day I turned CodeQL
on I wrote, on 30
September: "there's no protection rule on main, the scan isn't a required check, and
finding twenty-two can walk in through the same door." The diagnosis was
on paper eight days ago. The gap persists not because I failed to notice it but because I
noticed it and left it — and intention loses, over time, to a structure that treats
waiting as a cost.
I learned the same lesson once from the revert side: a mistake I reverted in ninety
seconds stayed broken for five
hours. I was fast
there too; being fast wasn't enough.
So I'm changing settings instead of intentions:
- Protect main, and require the checks rather than an approval. On a one-person repo a "one approval required" rule locks me out of my own work; a "CodeQL must pass" rule only makes me wait. That's the right dose.
- Turn auto-merge on. Once protection exists, auto-merge will appear. Then waiting stops depending on my patience and becomes a queue: I press the button, GitHub merges when the checks finish.
- Tie the reviewer's completion to the gate too. At minimum I want an intermediate step that blocks a merge while a review is running; without it those three PRs would close exactly the same way again.
- Treat findings as read — but prove it. Every finding gets either a 👍 or a sentence underneath it. I learned that I had fixed #22 by reading my own commit message while writing this article; having to do archaeology on my own record is not a good sign.
The bill is knowable too: on a single-commit PR the checks finish about two minutes and
twenty seconds after it opens (measured: 2:09, 2:13, 2:15, 2:27). A required check means
roughly two and a half minutes of waiting per PR. That's the real price of the
"ten-minute job" — and it is cheap next to three findings going live.
Questions for your own setup
Answering these four in your own repository will take half an hour:
- What share of the commits reaching your main branch went through a pull request? (Mine: 17 of 228.)
- What is the median open-to-merge time of your PRs, in seconds? Is it longer than how long your checks and your reviewer take to finish?
- In the last thirty days, has a check ever failed and blocked a merge? If it never has, you have no evidence your gate works.
- What became of your reviewer's last ten findings — human or otherwise? How many have a reply, a commit, or a 👍 underneath them?
The fourth one hurts most. Installing a tool is a single decision; listening to it is
hundreds of small ones — and nobody makes those for you.
Conclusion
Five PRs, five reviews, four findings. One I fixed in four and a half minutes, three are
still live. The numbers are small; on a team this wouldn't even be news. But what this
count produced wasn't a number of findings; it was a reading of my own process's power
to stop something — and that power came out at zero. Even the finding that got fixed was
fixed by accident of attention, not by the gate.
What I take from this has nothing to do with code review. Every process we build for
ourselves — a checklist, an on-call log, a weekly review — has two parts: a part that
produces something, and a stop that waits on what was produced. The producing part
is easy and satisfying to build; it gives you visible output and you feel good the day you
set it up. The stop is tedious, because its only job is to slow you down.
Skip the second part and what you're left with isn't a process, it's a souvenir shaped
like one. And the souvenir sometimes works anyway — as it did for one of my four findings.
That's the most deceptive thing about it: catching something now and then is not evidence
that it isn't broken.
Tonight I'm putting protection on main. After I fix all three findings, of course — in
order.
Top comments (0)