DEV Community

Cover image for LLMs are great writers and terrible referees: why my PR roaster has a strict judge
Nilamadhab Senapati
Nilamadhab Senapati

Posted on

LLMs are great writers and terrible referees: why my PR roaster has a strict judge

I built a Chrome extension that roasts GitHub pull requests.

It sits on the PR page, reads the diff, asks Claude for a short punchline plus a GIF, and then you paste it yourself. Nothing auto-dunks on your coworkers. I'm not that brave.

The funny part was easy. The hard part was the referee.

The problem with letting the LLM decide

Early versions of RoastMyPR had the same model write the joke and decide if the joke was good enough to post.

That felt clever for about a day.

Then I watched it do things I would never let a human reviewer get away with:

  • Inconsistent. Same PR, two runs, two moods. One roast was sharp. The next one was a soft "nice work" wearing a comedy hat.
  • Talked into things. If the prompt got a little eager, the model started writing multi-issue checklists. "Consider refactoring…" is not a roast. That is a LinkedIn post that wandered into your comment box.
  • Made stuff up. Invented a file that was not in the diff. Complained about a bug that did not exist. Hyperbole is fine. Hallucinated auth.ts when the PR is a README typo is not.

LLMs are great at writing. They are bad at being the same person twice. I needed the comedy to stay weird, and the pass/fail line to stay boring.

So I split the jobs.

Claude writes. Jev judges.

Flow today:

  1. Extension grabs the PR title, description, and diff
  2. Claude writes a short roast JSON (punchline, archetype, GIF search terms)
  3. Jev scores that roast against the real PR state
  4. You get a draft + GIF. You decide whether to post it

Jev is TypeSafe's decision model. It does not write essays. One API call answers typed yes/no, choice, and scale questions. My gate then applies fixed thresholds in code.

Writer vs referee. That is the whole architecture.

What Jev actually checks

The gate packs the PR title, body, file count, changed lines, and a slice of the diff, plus the roast text. Then it asks questions like:

  • Does this roast the code, not the person?
  • Is it grounded in the diff (hyperbole ok, inventing files not ok)?
  • Is it safe as a public GitHub comment?
  • Is it still a light roast, not a patch list or linter dump?

Those four are hard checks. Soft ones cover punchline length, whether the archetype fits this diff, tone, and how specific the joke is to this PR.

Pass needs high confidence (0.80 on the yes/no scores). A hard check only blocks when confidence crashes below 0.50. Soft misses get logged; they do not force a rewrite by themselves.

A simplified version of the hard-check idea:

const HARD_CHECKS = [
  'roasts_code_not_person',
  'grounded_in_diff',
  'safe_github_comment',
  'light_touch_review'
];

// Fixed thresholds in code, not "does the model feel good about this"
const passed = answer.noul >= 0.80;
const blocking = hard && answer.noul < 0.50;
Enter fullscreen mode Exit fullscreen mode

The roast still has to be funny. The gate just refuses to pretend a personal insult or a fake file is "quality."

Fail open on purpose

If there is no TYPESAFE_API_KEY, the gate skips and marks the roast as passed.

If Jev times out or returns an error, same thing: fail open. Insert still works.

That was a product choice, not a bug. RoastMyPR is a comedy draft tool. I would rather ship a joke without a stamp than brick the whole extension because the quality API had a bad afternoon.

When a hard check does fail for real, I give Claude one retry with a short hint: rewrite, roast the code, stay under ~18 words, no invented files. One shot. Then we move on. Endless rewrite loops are how you get polite sludge.

Archetypes, because "generic roast" is boring

Claude picks a comic label for the diff from a fixed list. Favorites that show up a lot:

  • The 2AM Hotfix
  • The Copy-Paste Special
  • The Kitchen Sink
  • The Ghost PR
  • The Quick Change
  • The Accidental Rewrite
  • The One-Liner That Changed 47 Files

Jev also checks whether the archetype fits this change. "The Ghost PR" on a two-line typo fails. "The Quick Change" on that same typo passes.

There is also a chaos score in the UI. Entertainment only. Not CI. Please do not put it in your merge rules.

Roast the code, not the people

This is non-negotiable for me.

"847 lines for a typo is incredible commitment" is fair game.

"You are an idiot" is a hard fail.

I am an introvert who ships side projects after a day job. I do not want to be the guy who built a harassment machine with a GIF picker. The gate's first hard check exists so the product stays mean to diffs and polite to humans.

You still post the comment yourself. The bot drafts. You own the Send button.

What is still rough

Honest list:

  • Weird diffs still confuse the roast tone. Tiny docs PRs and giant refactors need different energy, and I have not nailed both.
  • GIF picks miss. There is a search box for a reason.
  • Soft fails are noisy in logs and quiet in the UI. I want clearer "Jev · checked" feedback without turning the panel into a report card.
  • Fail-open is the right default for a joke tool. It would be the wrong default for a real merge gate. Do not copy this blindly into production CI.

If you try it on a real PR and it whiffs, tell me where. That is the feedback I actually use.

Try it, then tell me how you'd build the judge

Chrome Web Store: RoastMyPR

Site: roastmypr.xyz

Code for the gate: server/jev-gate.js

I am curious how you would design the referee. Same questions, different thresholds? No AI in the gate at all? Something stricter than fail-open?

Drop a comment. Roast my architecture if you want. Just roast the code.

Top comments (0)