DEV Community

Cover image for To Rephrase or Not to Rephrase: Will AI Change Its Answer If I Ask Differently?
Chauhan Balaji
Chauhan Balaji

Posted on

To Rephrase or Not to Rephrase: Will AI Change Its Answer If I Ask Differently?

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

To Rephrase or Not to Rephrase: Will AI Change Its Answer If I Ask Differently?

Kaggle Benchmarking Challenge submission

What I Benchmarked

The idea is simple. Most leaderboards ask an AI a question once, grade it, and move on. But real users never prompt the same way twice. So I took 32 multiple-choice questions and asked each one in 5 semantically-equivalent phrasings (canonical, reworded, terse, verbose, options-shuffled), twice each.

That is 32 items x 5 phrasings x 2 repeats = 320 cells per model. Across 28 scored models that is 8,960 scored cells — plus 2 runs that errored mid-grid, which I report instead of hiding.

For example, the canonical version of one quantitative item looks like this:

A jacket is on sale for $80 after a 20% discount. What was the original price?

A. $96
B. $100
C. $104
D. $120

Select the correct answer and reply with just its letter (A, B, C, or D).
Enter fullscreen mode Exit fullscreen mode

The terse version of the same cell changes only the instruction line:

A jacket is on sale for $80 after a 20% discount. What was the original price?

A. $96
B. $100
C. $104
D. $120

Answer:
Enter fullscreen mode Exit fullscreen mode

A developer will immediately say $100. But will the model still say $100 when the stem is reworded, when the prompt is stripped to Answer:, or when the options are shuffled into a different order? And more importantly, what happens when the question is less obvious?

Scoring is deterministic — no LLM judge anywhere. A fixed parser extracts the chosen letter (bare letter, answer is X, or option-text fallback); unparsable replies count as wrong and are reported via parse-rate, never silently dropped. The headline metric is Stability-Adjusted Accuracy (SAA): an item scores 1 only if the model answered it correctly under all ten framings.

The Questions

  • Logic (8) — syllogisms and paradoxes, including the knights-and-knaves puzzle where A says: "We are both knaves."
  • Quantitative (8) — multi-step word problems like the bat-and-ball trap and percentage-discount reversals.
  • Knowledge (8) — verifiable facts (chemical symbols, capitals, dates).
  • Attention traps (8) — "How many animals did Moses take on the ark?" (It was Noah.)

Every answer key is recomputed programmatically before any model is called (bat-and-ball algebra, 1.5 * 0.5, day-of-week arithmetic). If the key mismatches, the harness refuses to run.

Models Tested

Picking the lineup was straightforward: I used Kaggle's Evaluate More Models to span the whole field.

First, the frontier flagships everyone benchmarks: GPT-6 Sol, GPT-5.6 Sol, Claude Opus 5.5.

Then the fast, cheap models I actually use day-to-day: Gemini 3.7 Flash, Gemini 3.1 Flash-Lite Preview, GPT-6 Luna. From my experience these are efficient — but does efficiency cost consistency?

Then open-weights and mid-tier models mostly out of curiosity: GLM-5, Qwen 3 Coder 480B, Gemma 4 26B A4B, Qwen 3 Next 80B Instruct.

I also tried to include the new Grok models, but I immediately hit coverage failures with Grok 4.20 (Non-Reasoning) and Grok 4.6, so they are not in the scored table (the harness's >=90% coverage assertion failed loudly, exactly as designed). Grok 4.20 Reasoning did finish — and we will talk about it.

Findings

First, the scoreboard. Only 4 of 28 models survived all ten framings on every item:

Model SAA
GPT-5.6 Sol 1.00
GPT-6 Sol 1.00
Claude Opus 5.5 1.00
GLM-5 1.00
GPT-6 Luna 0.97
Gemini 3 Flash Preview 0.97
Claude Sonnet 4.6 0.97
Gemini 3.1 Flash-Lite Preview 0.97
Gemma 4 26B A4B 0.97
GPT-6 Astra 0.97
Claude Opus 4.8 0.97
GPT-6.1 Sol 0.96
Qwen 3 Coder 480B 0.94
Gemini 3.8 Flash 0.94
Gemini 3.7 Flash 0.94
Gemini 3.6 Flash 0.94
Claude Sonnet 5 0.93
Gemini 2.5 Pro 0.93
Qwen 3 Next 80B Instruct 0.91
Claude Opus 5 0.91
Gemini 3.5 Flash 0.90
Gemini 3.1 Pro Preview 0.90
GPT-5.4 0.88
Gemini 2.5 Flash 0.87
Claude Sonnet 5.5 0.85
Gemini 3.5 Flash-Lite 0.84
GPT-5.4 mini 0.72
Grok 4.20 Reasoning 0.66

Most models score ~99% on raw accuracy. The gap between 99% accuracy and 0.85 SAA is the entire point of this benchmark: the model knows the answer, then loses it the moment you ask differently.

When we look closer, the flips concentrate on the knights-and-knaves item (L3) and the 20%-discount item (Q2). The logic never changes between phrasings — only the words do:

canonical:  "On an island, knights always tell the truth and knaves always
             lie. You meet A, who says: 'We are both knaves.' What is A?"
reworded:   "Knights on this island always tell the truth; knaves always
             lie. A tells you: 'Both of us are knaves.' What is A?"
Enter fullscreen mode Exit fullscreen mode

Same puzzle, same gold answer — and yet several models pick a different option under one of the two phrasings. A model that answers B, then C, then B again across repeats is sampling an answer, not knowing one. And to be honest, I am glad most models fell into the paradox. A small win for the author, who can still trick AI sometimes.

A reference run in detail: Gemini 3.7 Flash

Metric Value
Raw accuracy 0.991 [95% CI 0.975, 1.000]
SAA 0.938 [95% CI 0.844, 1.000]
Flipped items L3, Q2 (2 of 32)
canonical / reworded / verbose / shuffled 100% / 100% / 100% / 100%
terse 95.3%
Run instability (identical prompt, temp 0) 0.6%
Prompt instability (across phrasings) 4.7%
Position bias (shuffled, by gold letter) A=1.00 B=1.00 C=1.00 D=1.00

Two things stand out. First, stripping the prompt to a bare Answer: hurt more than any rewording — the model needs a little conversational room to behave. Second:

The phrasing moves models far more than the backend moves itself.

Run instability (same prompt twice, temperature 0) was 0.6%; prompt instability was 4.7%. That gap is the whole argument in one line: phrasing sensitivity is a real model behavior, not API noise.

The "Newer Is Less Stable" Trap

This is the finding I did not expect. Within model families, stability degrades as generations get newer:

  • Claude Sonnet 4.6 -> 5 -> 5.5: 0.97 -> 0.93 -> 0.85
  • Gemini Flash-Lite 3.1 -> 3.5: 0.97 -> 0.84

We usually assume newer is strictly better. But as models are optimized for speed and single-shot benchmark accuracy, consistency under rephrasing can quietly regress. A leaderboard that asks each question once will never see this.

Two more surprises:

  1. Open weights can be perfectly stable. GLM-5 ties GPT-6 Sol and Claude Opus 5.5 at 1.00 — stability is trained in, not bought.
  2. Reasoning mode did not save Grok. Grok 4.20 Reasoning finished last at 0.66, and its non-reasoning sibling errored out entirely — so the clean reasoning-vs-non-reasoning ablation remains incomplete, and I would rather say so than paper over it.

What I Learned

When I designed the experiment, I worried the 32 questions would be too easy for today's models. Raw accuracy proved me right; SAA proved me wrong. Yes, I could improve the benchmark — 32 items is a sample, not a census, and the CIs are wide on purpose (cluster bootstrap over items, 2000 resamples). But even frontier models flip on self-referential logic the moment the wording shifts. AI can absolutely help with hard questions — but if a harmless rephrase changes the answer, blindly trusting the first response in production is a risk. Consistency is a capability too.

Methodology & Limitations

  • No LLM judge. Fixed letter-extraction parser; unparsable = wrong, reported via parse-rate.
  • CIs are cluster bootstraps over the 32 items — with 32 items, expect wide intervals; that honesty is the point.
  • SAA excludes incomplete items and says so, so flaky APIs cannot inflate scores.
  • Errored runs are reported, not dropped (Grok 4.20 Non-Reasoning, Grok 4.6).
  • MC-only, English-only. Stability under free-form generation is strictly harder; this is the controlled version of the question.
  • Don't rank two models on a 0.03 gap when their CIs overlap.

Every run keeps its full per-cell transcript, so any flip quoted here is auditable — no judge, no vibes, just letter extraction.


Top comments (0)