DEV Community

StareBrain
StareBrain

Posted on

How Do You Know Your Reviewer Is Still Reviewing?

We've been building a flag-for-a-human step into StareBrain for actions whose outcome comes back ambiguous. The obvious next move, once you have that step running for a while, is to automate the easy cases — write a rule that catches the pattern a human keeps resolving the same way, and let it handle those without a person clicking through each time.

That move has a failure mode nobody had named until a comment on our IH thread today. Automating the easy cases removes the person from the easy cases. Which means the only thing left checking whether the rule is still right is a person who no longer looks.

The actual risk

Say a rule graduates from "human confirms every time" to "auto-resolve unless a human objects." For a while this works. Then a near-miss shows up — a case that looks like the pattern the rule handles, but isn't. If nobody's really reading the suggestions anymore because the rule has been right for weeks, the near-miss goes through unchallenged. You don't find out from the review queue. You find out later, from whoever was affected by the wrong resolution.

The failure isn't "the rule was wrong." Rules are wrong sometimes; that's expected and survivable if someone's watching. The failure is "the rule was wrong, and nobody was positioned to notice."

The fix that came out of the thread

Don't jump straight from "human decides" to "rule decides." Add a middle stage: the rule drafts a resolution, a human still has to click to confirm it. You're not saving decision time yet, you're testing whether the rule's suggestions actually agree with what a human would have chosen. Count the disagreements over a couple of clean weeks. Only then does it earn full autonomy.

This part is fairly standard, shadow-mode before cutover. The sharper idea came next.

Testing the reviewer, not just the rule

Once a rule is running in suggest-only mode, a low disagreement rate is ambiguous by itself. It could mean the rule is good. It could also mean the reviewer stopped reading the suggestions and is clicking confirm on autopilot. Those look identical in the data: zero disagreements either way.

The fix: seed a known-wrong suggestion into the stream now and then, and watch whether the reviewer catches it.

Two details make this actually work, both of which I'd have gotten wrong on my own:

The seed has to be a plausible wrong, not an obvious one. An absurd error gets caught by anyone regardless of whether they're really paying attention, so it tells you nothing. The seed needs to look like a real near-miss, something wearing the correct label convincingly enough that catching it actually demonstrates attention.

Tell the reviewer canaries exist. My first instinct was that a secret seed is a more honest test, closer to a real error slipping through. That's backwards. The first time someone discovers an unannounced canary, they stop trusting every edge case that reaches them afterward, which is a worse outcome than losing some test purity. Announcing that canaries exist, without saying which specific ones, keeps the vigilance test intact without turning it into a trap.

Reading the result as a pair

The useful number isn't the seed catch rate alone. It's the seed catch rate next to the live disagreement rate, read together:

High catch rate, live disagreements still coming in → the rule is doing its job, the reviewer's doing theirs.
High catch rate, zero live disagreements → could genuinely mean the rule is very good.
Low catch rate, zero live disagreements → the reviewer has stopped looking, whether or not the rule is actually fine.

That third case is the one a raw disagreement count alone can't distinguish from "the rule is great." The seeded rate is what breaks the tie.

Where this leaves us

We haven't built any of this yet, it came out of a comment thread today. But it's specific enough to build directly: suggest-only stage, a small library of plausible wrong resolutions to seed at random intervals, a disclosed-but-unspecified canary policy, and a dashboard that shows both rates side by side instead of just one.

If you're already running something like this — automated resolution with a human check behind it — how do you verify the human check itself hasn't quietly become a rubber stamp?

Top comments (0)