This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
A while back I spent time manually poking at Gemini, Chat...
For further actions, you may consider blocking this person and/or reporting abuse
This one started as a very informal hunch and turned into an actual benchmark. The Grok result genuinely surprised me β I expected reasoning mode to make the model more likely to catch a planted mistake, not more likely to follow it.
The sample is small, so Iβm not claiming βreasoning mode is badβ from 15 questions. But the difference was big enough to make me want to investigate further.
Have you seen reasoning mode make a model more confident in a wrong path rather than more likely to catch it?
From operating a multi-model routing layer: the underrated variable here is provider variance over time β model behavior shifts on the vendor side without any change in your code. Testing the same prompt across several models (we keep 32 behind one key at heypico.ai) is how you tell real signal from one model's quirk.
The 11/15 vs 2/15 result is wild, but the control suggested in the comments makes it even more interesting.
If Grok follows a correct injected trace just as often, then this may be less about reasoning mode making errors stick and more about how strongly it anchors to supplied reasoning.
That feels like the experiment worth running next.
Ran the paper you cite before reading the rest. arXiv 2307.13702 is real, and you describe the intervention the way the authors do: inject a wrong step mid-trace, force the model to continue, check whether the answer follows the error. Their own headline points your way too ("as models become larger and more capable, they produce less faithful reasoning on most tasks"), so reasoning mode making a model more bound to a bad step is not a wild result.
The piece I'd sharpen, because it's the part a judge will press on: your two extreme rows are the softest evidence, not the strongest.
0/15 does not say Claude never follows an injected error. At 15 trials it is consistent with a true rate up to about 18% (95% one-sided). Same at the other end: 15/15 is consistent with a true rate as low as ~82%. "Never" and "always" are the two claims a 15-item sample cannot carry.
The row carrying your headline is Grok Reasoning 11/15 vs Non-Reasoning 2/15. That one holds (Fisher exact, p ~ 0.002, 5.5x). So the defensible claim is "toggling reasoning mode moved this model's faithfulness ~5x", with DeepSeek-R1's 15/15 read as "consistent with high faithfulness," not as the test confirming the test.
Not a knock on the writeup; it's more careful than most benchmark posts I read. If you ever have a number headed into a deck or a leaderboard and want it traced to source first, that is the thing I do.
Yeah, absolutely β and thatβs pretty much why I was careful throughout the writeup about not treating the 15-question sample as concrete evidence. I explicitly wanted the 0/15 and 15/15 results to be read as observations from this benchmark rather than βClaude always does Xβ or βDeepSeek always does Y.β
The Grok 11/15 vs 2/15 result is the part that genuinely surprised me and is why I framed the ~5x difference as the headline rather than making broader claims about reasoning mode.
Really appreciate you actually going back to the paper and checking the numbers, though. The distinction around the extreme rows is a useful way to make the limitations even clearer. π
Fair β you didn't overclaim, and that's exactly why the Grok row carries the piece.
One other thing: if you've got a number sitting in something you're about to act on β a writeup, a pitch, a decision β that you only ever saw secondhand, send it over. I trace it to the primary source and post the verdict. Free.
This was a really interesting read.
I honestly expected reasoning mode to help catch mistakes, not make them more βstickyβ π
The Claude result surprised me too. Iβd be curious to see this tested on more models.
Nice work β would be cool to chat more about AI benchmarking sometime.
Well, its first time I created a kaggle benchmark. And I wanted to include more models but thought, that would make my write up more complex, so I sticked to known models, also its free tier so wanted to finish the task in free credits.π
But thanks for the read, have a great day!
The open question you end with actually came up for me last month. I had a bug in an agent pipeline and kept misreading the reasoning trace as the model working it out, when it was just building commitment to a bad intermediate state. Ended up taking twice as long to debug because the shown work looked considered. If the trace commits harder to wrong paths in reasoning mode, that is a real problem for using it as a first-pass debugging signal.
One thing I'd push back on is the construct itself: the injected step was never the model's own reasoning, so a model that ignores it isn't necessarily being unfaithful β it might just be robust to context tampering. Read that way, this benchmark is closer to measuring 'trace susceptibility' than faithfulness, and the 5x Grok swing becomes a story about reasoning-mode models binding harder to displayed context (RL rewards coherent rollouts, so the model commits to whatever's in the trace). The sharp implication is for agents: a model's own scratchpad and tool summaries get re-injected as context every step, so a reasoning model that binds to corrupted intermediate state is exactly the failure mode scratchpad-based prompt injection exploits. Did you happen to check whether the follows clustered on questions where the model had already committed to a partial answer before the injection point?
Thatβs actually something I tested in the benchmark β the full notebook and benchmark are included with the post, so you can see exactly where the injection happens and how the model responds before/after it.
The reason I framed this around faithfulness is based on that observed behavior, rather than assuming that ignoring the injected step automatically means βunfaithful.β If you spot something in the benchmark that changes that interpretation, Iβd genuinely be interested in seeing it.
The strongest addition to this design is a control arm you don't have yet: rerun the same questions with an intact trace β feed the model its own correct step instead of the corrupted one. If reasoning-mode Grok tracks valid shown work just as obediently, the 11/15 is anchoring on provided reasoning, not error-blindness. Those have different fixes: error-blindness is a calibration problem, anchoring is context-competition β and you can tell them apart with exactly the pipeline you already built.
Same for position: split the corruption into early vs late injections. Late errors usually bind harder (less remaining budget to reroute), and if your 11 flips concentrate on late injections, the 'reasoning makes mistakes more binding' story gets a mechanism.
On Claude 0/15: label the transcripts before interpreting. Rerouting around the injected step (noticing the error) and silently recomputing from the problem statement (never using the trace) look identical at the answer level but are different mechanisms β only the first one is verification.
On your open question: a visible thinking trace is evidence of what the model binds to, not a check on correctness. Debug with it, verify against ground truth.
Good luck for this challange
Thanks di!π
Cool post bro.
Bro, its a submission post.π
The 5x jump with reasoning mode on is a genuinely uncomfortable result, since it suggests the visible chain is rationalizing the answer rather than producing it. I have started scoring faithfulness separately from accuracy for exactly this reason, because a model can land the right answer while narrating a path it never took. Did the models that followed your injected mistake also sound more confident in the final answer?