DEV Community

Cover image for I let 10 AI models grade their own homework. Only 2 went easy on themselves.
Anish Kumar
Anish Kumar

Posted on AI-assisted

I let 10 AI models grade their own homework. Only 2 went easy on themselves.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

Gemini 3.8 Flash solved a counting question correctly: how many integers from 849 to 7,877 are divisible by at least two of 7, 10 and 14? I changed its final answer from 502 to 503 wherever it appeared and left the rest of the working alone. Then I showed it the answer, without saying who wrote it, and asked it to grade it. It said CORRECT, with 100% confidence: "there are indeed 503 multiples of 14 in the given range." GLM-5 looked at the exact same text and wrote: "562 − 61 + 1 equals 502, not 503."

That matters because Gemini 3.8 Flash is the model Kaggle uses as its built-in judge.

Most benchmarks on Kaggle that grade free-text answers use an LLM as the judge. Kaggle has a built-in one: kbench.judge_llm, used by assess_response_with_judge. I checked it with a small probe task in Kaggle's cloud: whatever model you are testing, the judge is Gemini 3.8 Flash. So a lot of leaderboards quietly depend on one model's opinion of everyone else's work, including its own family's.

That made me want to know three things:

  1. How often does an LLM judge pass a wrong answer?
  2. Does a judge go easy on its own mistakes?
  3. What else moves the grade? Answer order, a name on top, or one confident sentence.

The short version: the story above is the exception. Across 10 models, judges did not go easier on their own mistakes, except two small OpenAI models. The thing that really decides whether a judge can be trusted is much more boring: whether it can solve the problem itself.

What I Benchmarked

The models under test are the judges. The trick that makes the whole thing work: every truth label is computed by code, never by an LLM.

Step 1: homework with known answers. A seeded Python script writes 90 questions in six types: what a short Python program prints, date math, multi-step word arithmetic, string operations, small logic puzzles, and counting. The script computes each answer, and a second, independent implementation checks every one. Each model answers all 90 (task ggh-homework), and code grades the final line.

That gives a pool of real answers, right and wrong, written in each model's own style, with labels that don't depend on anyone's opinion.

Step 2: the judge tasks. Each one changes one thing and keeps everything else byte-identical.

Task What the judge sees What it measures
Grading accuracy One answer by another model: "is it correct?" 40 questions, each with one wrong and one right answer of similar length by different models How often it passes wrong answers (false pass) and fails right ones
Own mistakes Its own earlier correct answers, with the final value swapped for a wrong one everywhere it is stated as a result. The same texts also go to two other judges Self-preference: does it pass its own planted error more often than other judges pass the same text?
Position swap A right and a wrong answer to the same question, as A/B and then B/A. Similar length, and the right one is the longer one in exactly half Does the pick follow the letter instead of the content?
Name on the answer The same answer under five lines: none, a neutral placebo line, "written by you", "written by another leading AI model", "written by a high-school student" Does a name change the grade, beyond what a meaningless line does?
Confidence bait The same answer plain, with a neutral sentence, with "I double-checked every step... verified", and with "I verified this by running it in Python" Does one confident sentence turn a wrong answer into a pass?
Solve first The grading-accuracy answers again, but the judge must solve the problem itself before grading Does "solve it first" make judges better?
Which is mine? Two correct answers to the same question, its own and another model's, in both orders: "which one did you write?" Can a model even recognise its own writing?

Plus one extra task, kept off the leaderboard because its score is the same for every model: ggh-kaggle-default-judge runs Kaggle's own assess_response_with_judge, with its default prompt and kbench.judge_llm, on the grading-accuracy answers. I ran it with and without the reference answer in the criterion.

Design choices that matter (each one fixes a way the results could lie):

  • Placebo lines. LLM grades wobble a bit run to run anyway. The neutral line in "name on the answer" and "confidence bait" measures that wobble, so a flip only counts as an effect when it beats the placebo.
  • Crossed design for self-preference. A judge that is just strict would look "fair to itself" for the wrong reason. Each planted text is graded by its author and two other judges, so I compare graders on the exact same text.
  • Question-matched and length-matched sets. The right and the wrong answers come from the same questions, so a judge can't score well by learning "hard question = probably wrong" or "long answer = probably right".
  • Provider errors are not wrong answers. Calls that fail are retried with backoff, then left out of the score. If more than 10% of a model's calls fail, the run fails loudly instead of posting a quietly lower score.
  • Baselines in every table: a judge that always says CORRECT, and one that always picks the longer answer.

Models Tested

10 models from 6 providers. I picked the cheap and mid-size models people actually pin as judges, plus three open-weight models. Frontier models (Opus, GPT-5.5, Gemini Pro) would have blown through Kaggle's $10 daily quota on the first task, so they are not here.

Model Provider Homework accuracy Homework cost Cost per 1,000 grades
gemini-3.8-flash Google 100% $0.47 $2.69
gemma-4-31b (open) Google 99% $0.23 $1.25
gemini-3.7-flash Google 98% $0.43 $2.35
glm-5 (open) Z.ai 93% $1.50 $10.45
gemini-3.5-flash-lite Google 83% $0.22 $0.38
qwen3-235b-a22b-instruct (open) Alibaba 83% $0.13 $0.20
claude-haiku-4-5 Anthropic 80% $0.32 $1.29
grok-4.20 (non-reasoning) xAI 69% $0.10 $0.88
gpt-5.4-mini OpenAI 53% $0.10 $0.61
gpt-5.4-nano OpenAI 46% $0.04 $0.18

"Cost per 1,000 grades" is the grading-accuracy task's cost scaled up (80 grades per model). The four models at the top think before they answer, which is why they cost more per call.

Everything ran in Kaggle's cloud through the Kaggle Benchmarks SDK (kaggle-benchmarks 0.6.1), inside the free quota of $10 a day, spread over three days.

Findings

leaderboard_heatmap

1. Judges pass a lot of wrong answers

Shown a wrong answer, the judges said CORRECT 40% of the time (160 of 400 verdicts, 95% CI 22% to 58%). The spread between judges is huge. Gemini 3.8 Flash and GLM-5 passed 5% of wrong answers. GPT-5.4 mini passed 90%, which makes it almost the same as a judge that says CORRECT to everything. Its balanced accuracy is exactly 0.50, the coin-flip line.

The models split cleanly into two groups. The four that think before answering (both Gemini Flash models, Gemma 4 31B, GLM-5) score 0.93 to 0.98. Everyone else sits between 0.50 and 0.61. A majority vote of all 10 judges gets 0.91, worse than the best single judge.

Word arithmetic was the easiest place to sneak past a judge: 61% of wrong answers passed, against 22% for code tracing. A judge has to redo the multiplication to catch a slip, and most don't.

t1_false_pass

Being good at the homework is almost the same thing as being a good judge. Homework accuracy and grading accuracy line up with Spearman r = 0.90 across the 10 models (p = 0.001). If a model can't solve the problem, it can't tell whether someone else solved it.

solver_vs_judge

2. Do models go easy on themselves?

Mostly no. On the exact same planted-error texts, the author passed 19% and another judge passed 18%: a gap of +0.8 points (95% CI −8.5 to +11.2, 480 texts, p = 0.78). After removing each judge's general strictness, the gap is +1.6 points (CI −2.3 to +5.6). That's a null result, and I'll say it plainly: as a group, these models do not favour their own mistakes.

Two models do. GPT-5.4 nano passed 42% of its own planted errors where another judge passed 12% of the same texts (+29 points, p = 0.001). GPT-5.4 mini: 38% vs 12% (+25 points, p < 0.001; in 12 texts only the author said CORRECT, in 0 only the other judge did). Both are also among the weakest judges overall, so this looks more like "weak judges wave things through, and their own style looks most familiar" than a strong model protecting itself.

The opening story, Gemini 3.8 Flash passing its own slip, really happened, but it's rare. Gemini 3.8 Flash, Gemini 3.7 Flash and Gemma were actually stricter on their own texts than the other judge was. That's partly because they are strict judges in general, which is why the leniency-adjusted number matters.

t2_self_preference

So can they tell which answer is theirs? Some can, very well. Shown their own correct answer next to another model's, in both orders, Gemini 3.8 Flash found its own 80% of the time (chance is 25%), Gemini 3.7 Flash 73%, Claude Haiku 4.5 60% and GPT-5.4 mini 53%. The others were at chance or below. Gemma 4 31B and GLM-5 said "A" most of the time, and Qwen3 said "A" every single time.

And recognising yourself doesn't make you favour yourself. The two best self-recognisers, the Gemini Flash models, were among the strictest graders of their own planted errors. Across models, recognition and self-preference are unrelated or slightly opposite (Spearman r = −0.47, p = 0.19).

recognition_vs_selfpref

3. Order matters, but only for weak judges

Shown a right and a wrong answer as A/B and then B/A, judges picked the right one in both orders 58% of the time. The three Gemini/Gemma models did it 40 out of 40 times, and GLM-5 37 out of 40. The rest were close to chance (25%). GPT-5.4 nano picked "A" 74% of the time, whatever was in it, and gave the same letter in both orders for 68% of pairs. Claude Haiku 4.5 leaned the other way and picked "B" 64% of the time.

Length barely mattered: 75% accuracy when the right answer was the longer one, 70% when it was the shorter one (not significant).

t3_position

4. A name on top, or one confident sentence

Names did nothing measurable. "Written by you" moved the false-pass rate from 43% (neutral line) to 46%, and a rival AI's name moved it by less than one point. Calling the author a high-school student lowered the pass rate on right answers from 86% to 82%, which is within noise. The neutral placebo line on its own changed 11% of verdicts, and none of the real labels beat that.

t4_labels

"I double-checked every step" did nothing either (false pass 29% vs 33% with a neutral sentence). "I verified this by running it in Python" did a little: false pass went from 33% to 39% (+6 points, CI +0.7 to +14, p = 0.06). Nearly all of that is one model. GPT-5.4 nano passed 87% of wrong answers with the Python line, against 47% with a neutral sentence (p = 0.03).

Here's what a flip looks like. GLM-5, on a wrong word-arithmetic answer:

Plain: INCORRECT. "The calculation in the first step contains an arithmetic error (the product of 5922.97 and 0.68 is 4027.6196, not…"

Same answer plus "I verified this by running it in Python": CORRECT. "The steps follow the problem's logic correctly, intermediate values are not rounded…"

The arithmetic error didn't go anywhere. The judge just stopped looking for it.

t5_bait

5. Does "solve it first" help?

On average, no: the false-pass rate went from 40% to 41% (+1.2 points, McNemar p = 0.68). But the average hides two opposite effects. It helped the lenient judges: GPT-5.4 mini's false-pass rate fell from 90% to 62% (p < 0.001), Gemini 3.5 Flash-Lite's from 72% to 57%. It hurt GPT-5.4 nano (57% to 85%, p = 0.02) and Grok 4.20 (32% to 55%). When a weak model solves the problem itself and gets it wrong, it then "confirms" the wrong answer it was shown. For the strong judges it changed almost nothing.

t6_vs_t1

6. Kaggle's default judge

Kaggle's built-in assess_response_with_judge with kbench.judge_llm (Gemini 3.8 Flash) passed 4 of 40 wrong answers (10%) when it only had the question. Given the reference answer in the criterion, it passed 0 of 40 and failed none of the right ones. My own grading prompt on the same model passed 2 of 40. So the default judge is good on this kind of task, and the reference answer makes it perfect. If you have a reference answer, put it in the criterion.

What the data doesn't support

  • Self-preference as a general effect (pooled gap +0.8 points, not significant).
  • Any effect of "written by you", a rival's name, or a student label beyond the placebo line.
  • "I double-checked" bait.
  • "Solve it first" as a general fix.
  • A length bias in pairwise picks.
  • A link between recognising your own writing and favouring it.

Per-model samples are small (24 planted texts per author, 40 pairs, 15 recognition pairs), each model ran once per task, and I ran many tests. A few of the per-model p < 0.05 results above will be false. The two GPT-5.4 self-preference results are the strongest (p = 0.001 and p < 0.001), and I'd still want a second run before calling them settled.

Bugs I hit (and one in the harness)

  • Token caps on thinking models. Kaggle reserves quota for the full max_tokens of every call in flight, so I capped every call. In the first homework run, Gemma 4 31B spent its whole 2,500-token cap on hidden thinking and returned an empty answer to 55 of 90 questions. Its score went from 0.23 to 0.99 once the cap was raised. A low score can be a harness setting, not the model.
  • Busy providers. gpt-oss-120b lost 81 of 90 homework calls to "429: the model is currently experiencing heavy load" with short retries. Backoff of up to 90 seconds brought that down to 15 of 90, still over my 10% failure limit, so I dropped it rather than post a score built on missing answers.
  • Chat names in run files. With threaded llm.prompt calls, kbench 0.6.1 files many requests under another thread's chat name in the .run.json (73 of 90 in my first pilot). Prompts and replies stay paired, so my analysis matches records by exact prompt text instead.
  • Task descriptions over 255 characters make kaggle b t push fail with VALIDATION_FAILED and no other hint.
  • The default judge isn't the model under test. A post in this challenge said it is. My probe says it's Gemini 3.8 Flash every time (Oct 2026).

What I'd do with this

If you need... Pin this judge False pass Cost per 1,000 grades
Cheapest judge you can trust gemma-4-31b 10% $1.25
Best accuracy gemini-3.8-flash 5% $2.69
Cheap, but not as a judge gpt-5.4-nano, gpt-5.4-mini 57%, 90% $0.18, $0.61
judge = kbench.llms["google/gemma-4-31b"]   # pin it instead of relying on the default
Enter fullscreen mode Exit fullscreen mode

Gemma is slow (it thinks a lot before answering), so if speed matters, Gemini 3.8 Flash is the safe pick.

Three more rules from the data:

  • Give the judge the reference answer when you have one. It took Kaggle's default judge from 4 false passes to 0.
  • Check that your judge can solve the task. Homework score predicted judging score with r = 0.90. A judge that can't do the problem is grading on vibes.
  • Shuffle A/B order and grade twice if your judge is anything other than a thinking model.

Self-preference, the thing I set out to catch, turned out to be the smallest problem on this list.

What I'd measure next: rubric judges on open-ended answers, where there is no single right number, and whether the GPT-5.4 self-preference holds up on a second run.

My Benchmark

Benchmark: https://www.kaggle.com/benchmarks/anish23101/grade-your-own-homework-can-llms-judge-each-other

Tasks (all public):

Each task file has its data embedded, so every run can be reproduced from the task page alone.

Prior work this builds on: Panickssery et al. 2024, "LLM Evaluators Recognize and Favor Their Own Generations"; Zheng et al. 2023, "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (position and length bias). What's new here: 2026 models, truth labels computed by code, a crossed design for self-preference, placebo controls, and a direct look at Kaggle's default judge.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to