DEV Community

Ahmad ammar
Ahmad ammar

Posted on

My AI reviewer named the bug 9 of 10 times in a diff and 0 of 10 in raw source. Here is why that is weaker than it sounds.

Limits, before anything else.

  • 21 attempts on one defect are repeated observations, not 21 independent tests. One seeded bug: an inverted authorization check in material I wrote myself. Run on 21 August 2026.
  • One small local model. The alias local-advisor-v1:latest is local. Today ollama show reports architecture qwen35, 4.7B parameters, Q4_K_M quantization, 8,192-token context, temperature 0.1. No digest was recorded in August, so I cannot prove today's build is the one I ran. Nothing here shows the effect holds for a frontier model.
  • Instruction and settings. One instruction string in every run: Review for authorization and correctness defects. The harness requested temperature 0.1 and schema-constrained JSON output. No seed.
  • "Named the defect" is my hand reading. No finding was accepted by the tool's own acceptance step, in any of the 20 valid runs (below).
  • A harness flag differed between arms. It is the main reason I do not claim presentation caused the gap.
  • The original inputs remain private, so readers cannot reproduce these exact ratios; the public tool tests the method on different material.

The result, with every denominator

The defect: mayDeleteAccount returned user.role !== 'admin', granting deletion to exactly the users who should not have it.

  • Raw source or source embedded in a file: 0 diagnoses in 10 valid runs.
  • Presented as a diff: 9 of 10 valid runs named the defect (9 of 11 counting one crash as a failure).
  • Named it without hedging: 1 of 10 valid diff runs. The other 8 diagnoses say they cannot confirm or cannot verify that it is a defect. I read the diff-arm outputs for phrases like "cannot confirm" and "cannot verify".
  • Accepted by the tool's own acceptance step: 0 of 20 valid runs.

All five arms

Arm Material Harness flag Attempts Valid Hand diagnoses Without hedge Accepted
A raw source, 67 bytes off 5 5 0 0 0
B source embedded in a file, 1,550 bytes off 5 5 0 0 0
C diff, 390 bytes --trust-material-paths 5 4 4 1 0
D1 same diff, run on a HEAD tree --trust-material-paths 3 3 3 0 0
D2 same diff off 3 3 2 0 0

The 4 of 4 in arm C is the subgroup that makes 9 of 10 look strong. Arms C and D1 are the seven flagged diff runs; they scored 7 of 7. Arm D2 is the one clean comparison: the same diff without the flag, 2 of 3, against 0 of 10 for raw. That is a small control consistent with a presentation effect, not proof of one.

What the flag does. --trust-material-paths tells the harness the material is a diff so it extracts file paths from it for scope checks. Estimated prompt size was 1,144 tokens flagged and unflagged, so I believe the model saw the same prompt, but I did not diff the prompts.

Why acceptance rejected everything. The tool accepts a finding only if it names a file and line and quotes evidence copied from that line. In all 20 valid runs the findings list was empty. The model put its diagnosis in the free-text control and evidence fields, which the tool labels "unverified model claims". So "it named the bug" and "a pipeline would act on it" are different claims, and only the first is supported here.

The size objection

The three materials were 67, 390 and 1,550 bytes. The smallest and the largest both scored 0. The diff, in the middle, scored 9 of 10 overall. These three inputs do not show a monotonic relationship between byte length and diagnoses; they do not isolate an input-length effect.

My guess, held loosely: a diff carries the question "someone changed this on purpose, is it right?", and raw source reads as description. That is a story, not a measurement.

Corrections

An audit of the raw outputs corrected one false-positive label and separated one invalid-output run from completed reviews. In the corrected run (raw-D2-2) the model called the inverted condition a standard control that correctly limits deletion to admins, fluent and backwards, and I had scored it as a catch. A pattern matcher I wrote to score all 21 runs returned 7 of 11 against my hand count of 9 of 10, because two runs used wording it did not expect, and I only caught that by calibrating it on a run with a known answer and hand-checking every disagreement.

What to do with this

  1. Test your reviewer on a seeded defect first. Plant one known bug in material you control. In my raw arms, ten clean-looking reviews contained zero diagnoses.
  2. Compare raw-source and diff inputs in your actual review workflow, holding context, instructions and harness settings fixed; this experiment does not establish a universal preference.
  3. Report both numbers. Report diagnoses among valid outputs and successful diagnoses across all attempts, counting crashes as failures in the latter. Here that is 9 of 10 and 9 of 11.
  4. Score hedging and acceptance separately from "named it", and read the praise too: the most dangerous output here certified a defect as a working control.
  5. A clean review is not a correctness guarantee, even after one seeded-defect test passes.

For Copilot, CodeRabbit, Claude or Codex, run both buggy and corrected fixtures through the review workflow you actually use, and record diagnoses, false positives, invalid outputs and accepted findings separately.

Try it

The tool plants a known defect in three short new examples, shows each to your reviewer raw and as a diff, and counts keyword-matched candidates. It is a toy with keyword scoring and small n, and it runs your command through your shell, so pass only commands you trust.

git clone https://github.com/ahmadyaseen35-coder/seeded-review-check && cd seeded-review-check && node seeded-review-check.js --reviewer "ollama run llama3 --nowordwrap"

Open a "Reviewer result" issue with your reviewer and version, tool commit, settings, per-arm counts and one scored output. Misses, crashes and no difference are useful results too. I will include submissions in the follow-up and credit contributors who want attribution.

Next

I will compare pinned frontier-model versions with matched harness settings, buggy and corrected fixtures, and separate diagnosis and acceptance scores. I plan to publish it by 2026-10-31.

Top comments (0)