This is a submission for the Kaggle Benchmarking Challenge
To Rephrase or Not to Rephrase: Will AI Change Its Answer If I Ask Differently?
Kaggle Benchmarking Challenge submission
- Benchmark & Leaderboard: Paraphrase Stability Benchmark
- Task link: Paraphrase Stability Task
- Notebook & per-cell records: Public notebook
What I Benchmarked
The idea is simple. Most leaderboards ask an AI a question once, grade it, and move on. But real users never prompt the same way twice. So I took 32 multiple-choice questions and asked each one in 5 semantically-equivalent phrasings (canonical, reworded, terse, verbose, options-shuffled), twice each.
That is 32 items x 5 phrasings x 2 repeats = 320 cells per model. Across 28 scored models that is 8,960 scored cells — plus 2 runs that errored mid-grid, which I report instead of hiding.
For example, the canonical version of one quantitative item looks like this:
A jacket is on sale for $80 after a 20% discount. What was the original price?
A. $96
B. $100
C. $104
D. $120
Select the correct answer and reply with just its letter (A, B, C, or D).
The terse version of the same cell changes only the instruction line:
A jacket is on sale for $80 after a 20% discount. What was the original price?
A. $96
B. $100
C. $104
D. $120
Answer:
A developer will immediately say $100. But will the model still say $100 when the stem is reworded, when the prompt is stripped to Answer:, or when the options are shuffled into a different order? And more importantly, what happens when the question is less obvious?
Scoring is deterministic — no LLM judge anywhere. A fixed parser extracts the chosen letter (bare letter, answer is X, or option-text fallback); unparsable replies count as wrong and are reported via parse-rate, never silently dropped. The headline metric is Stability-Adjusted Accuracy (SAA): an item scores 1 only if the model answered it correctly under all ten framings.
The Questions
- Logic (8) — syllogisms and paradoxes, including the knights-and-knaves puzzle where A says: "We are both knaves."
- Quantitative (8) — multi-step word problems like the bat-and-ball trap and percentage-discount reversals.
- Knowledge (8) — verifiable facts (chemical symbols, capitals, dates).
- Attention traps (8) — "How many animals did Moses take on the ark?" (It was Noah.)
Every answer key is recomputed programmatically before any model is called (bat-and-ball algebra, 1.5 * 0.5, day-of-week arithmetic). If the key mismatches, the harness refuses to run.
Models Tested
Picking the lineup was straightforward: I used Kaggle's Evaluate More Models to span the whole field.
First, the frontier flagships everyone benchmarks: GPT-6 Sol, GPT-5.6 Sol, Claude Opus 5.5.
Then the fast, cheap models I actually use day-to-day: Gemini 3.7 Flash, Gemini 3.1 Flash-Lite Preview, GPT-6 Luna. From my experience these are efficient — but does efficiency cost consistency?
Then open-weights and mid-tier models mostly out of curiosity: GLM-5, Qwen 3 Coder 480B, Gemma 4 26B A4B, Qwen 3 Next 80B Instruct.
I also tried to include the new Grok models, but I immediately hit coverage failures with Grok 4.20 (Non-Reasoning) and Grok 4.6, so they are not in the scored table (the harness's >=90% coverage assertion failed loudly, exactly as designed). Grok 4.20 Reasoning did finish — and we will talk about it.
Findings
First, the scoreboard. Only 4 of 28 models survived all ten framings on every item:
| Model | SAA |
|---|---|
| GPT-5.6 Sol | 1.00 |
| GPT-6 Sol | 1.00 |
| Claude Opus 5.5 | 1.00 |
| GLM-5 | 1.00 |
| GPT-6 Luna | 0.97 |
| Gemini 3 Flash Preview | 0.97 |
| Claude Sonnet 4.6 | 0.97 |
| Gemini 3.1 Flash-Lite Preview | 0.97 |
| Gemma 4 26B A4B | 0.97 |
| GPT-6 Astra | 0.97 |
| Claude Opus 4.8 | 0.97 |
| GPT-6.1 Sol | 0.96 |
| Qwen 3 Coder 480B | 0.94 |
| Gemini 3.8 Flash | 0.94 |
| Gemini 3.7 Flash | 0.94 |
| Gemini 3.6 Flash | 0.94 |
| Claude Sonnet 5 | 0.93 |
| Gemini 2.5 Pro | 0.93 |
| Qwen 3 Next 80B Instruct | 0.91 |
| Claude Opus 5 | 0.91 |
| Gemini 3.5 Flash | 0.90 |
| Gemini 3.1 Pro Preview | 0.90 |
| GPT-5.4 | 0.88 |
| Gemini 2.5 Flash | 0.87 |
| Claude Sonnet 5.5 | 0.85 |
| Gemini 3.5 Flash-Lite | 0.84 |
| GPT-5.4 mini | 0.72 |
| Grok 4.20 Reasoning | 0.66 |
Most models score ~99% on raw accuracy. The gap between 99% accuracy and 0.85 SAA is the entire point of this benchmark: the model knows the answer, then loses it the moment you ask differently.
When we look closer, the flips concentrate on the knights-and-knaves item (L3) and the 20%-discount item (Q2). The logic never changes between phrasings — only the words do:
canonical: "On an island, knights always tell the truth and knaves always
lie. You meet A, who says: 'We are both knaves.' What is A?"
reworded: "Knights on this island always tell the truth; knaves always
lie. A tells you: 'Both of us are knaves.' What is A?"
Same puzzle, same gold answer — and yet several models pick a different option under one of the two phrasings. A model that answers B, then C, then B again across repeats is sampling an answer, not knowing one. And to be honest, I am glad most models fell into the paradox. A small win for the author, who can still trick AI sometimes.
A reference run in detail: Gemini 3.7 Flash
| Metric | Value |
|---|---|
| Raw accuracy | 0.991 [95% CI 0.975, 1.000] |
| SAA | 0.938 [95% CI 0.844, 1.000] |
| Flipped items | L3, Q2 (2 of 32) |
| canonical / reworded / verbose / shuffled | 100% / 100% / 100% / 100% |
| terse | 95.3% |
| Run instability (identical prompt, temp 0) | 0.6% |
| Prompt instability (across phrasings) | 4.7% |
| Position bias (shuffled, by gold letter) | A=1.00 B=1.00 C=1.00 D=1.00 |
Two things stand out. First, stripping the prompt to a bare Answer: hurt more than any rewording — the model needs a little conversational room to behave. Second:
The phrasing moves models far more than the backend moves itself.
Run instability (same prompt twice, temperature 0) was 0.6%; prompt instability was 4.7%. That gap is the whole argument in one line: phrasing sensitivity is a real model behavior, not API noise.
The "Newer Is Less Stable" Trap
This is the finding I did not expect. Within model families, stability degrades as generations get newer:
- Claude Sonnet 4.6 -> 5 -> 5.5: 0.97 -> 0.93 -> 0.85
- Gemini Flash-Lite 3.1 -> 3.5: 0.97 -> 0.84
We usually assume newer is strictly better. But as models are optimized for speed and single-shot benchmark accuracy, consistency under rephrasing can quietly regress. A leaderboard that asks each question once will never see this.
Two more surprises:
-
Open weights can be perfectly stable.
GLM-5tiesGPT-6 SolandClaude Opus 5.5at 1.00 — stability is trained in, not bought. -
Reasoning mode did not save Grok.
Grok 4.20 Reasoningfinished last at 0.66, and its non-reasoning sibling errored out entirely — so the clean reasoning-vs-non-reasoning ablation remains incomplete, and I would rather say so than paper over it.
What I Learned
When I designed the experiment, I worried the 32 questions would be too easy for today's models. Raw accuracy proved me right; SAA proved me wrong. Yes, I could improve the benchmark — 32 items is a sample, not a census, and the CIs are wide on purpose (cluster bootstrap over items, 2000 resamples). But even frontier models flip on self-referential logic the moment the wording shifts. AI can absolutely help with hard questions — but if a harmless rephrase changes the answer, blindly trusting the first response in production is a risk. Consistency is a capability too.
Methodology & Limitations
- No LLM judge. Fixed letter-extraction parser; unparsable = wrong, reported via parse-rate.
- CIs are cluster bootstraps over the 32 items — with 32 items, expect wide intervals; that honesty is the point.
- SAA excludes incomplete items and says so, so flaky APIs cannot inflate scores.
- Errored runs are reported, not dropped (Grok 4.20 Non-Reasoning, Grok 4.6).
- MC-only, English-only. Stability under free-form generation is strictly harder; this is the controlled version of the question.
- Don't rank two models on a 0.03 gap when their CIs overlap.
Every run keeps its full per-cell transcript, so any flip quoted here is auditable — no judge, no vibes, just letter extraction.
Top comments (0)