DEV Community

Cover image for Rejudge Replaces Self-Review With 3 Independent Models and a Judge
Khasky
Khasky

Posted on

Rejudge Replaces Self-Review With 3 Independent Models and a Judge

I have a hard time trusting an AI coding agent to review code it just wrote.

The model already made the decision that the implementation was good enough to produce.

Then we ask it to inspect that same implementation using many of the same learned habits and assumptions.

Sometimes it catches mistakes. Sometimes it revalidates them. A fresh session helps. A different model helps more.

But once different models disagree, somebody still has to decide which review to trust.

Rejudge makes that disagreement-resolution step explicit.

The architecture

The same review request goes to three reviewers:

reviewer A
reviewer B
reviewer C
Enter fullscreen mode Exit fullscreen mode

Each works in an isolated context. They do not see each other's reasoning, tools, or conclusions. After all three finish, a separate judge receives their reports. The judge can ask follow-up questions when the panel disagrees. Then it writes one final answer.

same question
    |
    +--> reviewer A
    +--> reviewer B
    +--> reviewer C
             |
           judge
             |
        final answer
Enter fullscreen mode Exit fullscreen mode

Independence happens before collaboration.

How reviewers inspect the code

By default, reviewer tools include:

read
grep
find
ls
git_diff
web_search (if the host provides one)
Enter fullscreen mode Exit fullscreen mode

The normal reviewers do not get:

edit
write
bash
Enter fullscreen mode Exit fullscreen mode

That is a sensible default for code review. 🔒

How judge makes a decision

The judge gets no workspace access. It sees the three reports.

If those reports conflict, it can call ask_panel and request clarification.

That means the judge is not a hidden fourth reviewer. Its role is:

compare
challenge
adjudicate
synthesize
Enter fullscreen mode Exit fullscreen mode

Practical usage

Install (it needs Node.js 22.19.0 or newer):

npm install -g rejudge
Enter fullscreen mode Exit fullscreen mode

Review a diff:

git diff | rejudge "review this change"
Enter fullscreen mode Exit fullscreen mode

Ask a targeted question:

rejudge "does this migration need a lock?"
Enter fullscreen mode Exit fullscreen mode

Resume later:

rejudge --resume <run-id> "what about the rollback path?"
Enter fullscreen mode Exit fullscreen mode

The answer goes to stdout, while progress, the config in use and the run ID go to stderr, so redirecting stdout to a file leaves only the answer. A resumed run reopens the same sessions, and the new question goes to the judge first. The reviewers hear it only if the judge calls ask_panel.

Coding-agent integration

Rejudge supports:

  • CLI
  • native Pi tool
  • Agent Skill for coding agents outside Pi

Agent Skill install:

npx skills add syabro/rejudge -g -y
Enter fullscreen mode Exit fullscreen mode

The skills are a separate copy, so the README says to refresh them after each Rejudge release with npx skills update -g -y. Inside Pi, the extension is one more line:

pi install "$(npm root -g)/rejudge"
Enter fullscreen mode Exit fullscreen mode

Rejudge runs on Pi and reads its provider settings, so a key Pi already accepts works here too.

The models are configurable

Rejudge is not tied to one fixed provider combination. The config has a reviewer list and a separate judge model, with two reviewers as the minimum. Each model also gets a reasoning level from:

minimal
low
medium
high
xhigh
Enter fullscreen mode Exit fullscreen mode

The global file lives at ~/.config/rejudge/config.json, and a .rejudge/config.json in the project wins over it. That makes it possible to build a genuinely mixed panel.

There is an unsafe mode

--unsafe / --full gives reviewers:

edit
write
bash
Enter fullscreen mode Exit fullscreen mode

The docs explicitly say this is not a sandbox. The judge still gets only ask_panel. For review-only work, I would stay read-only.

Multiple providers mean multiple privacy boundaries

Every reviewer model sees the request. Anything a reviewer reads becomes part of that provider session.

The judge does not inspect the workspace directly, but reviewer reports may quote code.

Read-only tools stop local changes. They do not keep file contents private, and the README warns that instructions hidden in a request or in a file a reviewer opens can steer what it reads and reports.

Runs also leave records. Sessions are written to ${TMPDIR}/rejudge/runs/<run-id>/ while a run executes, and cleanup after roughly 24 hours is best-effort. The debugLog option is off by default, and turning it on writes full model thinking to .rejudge/logs/.

For proprietary repositories, this needs to be acceptable before running the panel.

The cost

A fresh review starts with:

3 reviewer calls
+ 1 judge call
Enter fullscreen mode Exit fullscreen mode

Then add tool loops, retries, recovery, and judge follow-ups. Rejudge has no spending cap of its own. This is not a free accuracy multiplier. It is a compute tradeoff. 💸

Where I would use it

I would reserve it for changes where mistakes are expensive:

database migrations
locking/concurrency
authentication
authorization
permissions
security-sensitive code
data deletion
rollback logic
complex refactors
Enter fullscreen mode Exit fullscreen mode

Rejudge is a structured independent second opinion. For consequential code, that can be worth the extra cost.

References


Follow me for more on AI and Software Development:

khasky — LinkedIn / Patreon / GitHub / Bluesky / Mastodon

khaskydev — X / Threads / Instagram / Pinterest / Facebook

Top comments (0)