DEV Community

Khushi .
Khushi .

Posted on

LLMs Pass the Data-Science Quiz, Then Give Different Advice: A Kaggle Benchmark of 36 Measured Judgment Calls

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

I spend a lot of time in Kaggle tabular competitions, and the decisions that cost me the most were never about model architecture. They were judgment calls: is this +0.0001 real? Should I append the original dataset? Which two submissions do I pick on the last day? (I once let the platform auto-pick, and an honest run that would have placed well ended at rank 708 of 3,575.)

So I built a benchmark where every answer was measured rather than taken from a textbook. Each case is a short competition situation whose correct call I established on Kaggle Playground S6E9 and S6E10 (Sep–Oct 2026), or that is a mathematical fact about the metric.

12 topics × 3 numeric variants = 36 cases:

# Topic What was measured
T01 p-values at large n 19/21 features had p < 1e-300, yet univariate AUC ranged 0.840 → 0.506
T02 Pearson vs Spearman with zeros delay columns: r = 0.84, ρ = 0.47, 91% exact zeros in both
T03 "0" that means not applicable wifi rating: satisfied rate 0.887 at 0, 0.391 at 1, 0.950 at 5
T04 AUC vs mutual information one rating ranks 10/21 by AUC but 3/21 by MI and Cramér's V
T05 CV vs leaderboard resolution public LB standard error ≈ 0.0003; +0.00005 CV gains never showed
T06 using the original dataset appending rows: −0.00025; model-on-original as a feature: +0.0007
T07 the more-folds trap 10-fold: OOF +0.00012, leaderboard +0.0000
T08 leak hunting 0 duplicates, 0 train/test overlap; 17 of 129,880 original rows shared
T09 calibration vs AUC AUC is rank-based, so a monotone recalibration cannot raise it
T10 blend diversity GBMs correlate 0.9985–0.9994; a weaker NN adds more to the stack
T11 final-submission selection auto-picked best-public files → rank 708/3575
T12 a step-shaped feature target jumps at age 39 and drops at 61

There are two tasks over the same 36 situations, each scored as cases correct out of 36:

  1. ds-judgment (recognise): four options. One is the measured answer, one is the classic trap (the rule of thumb my measurements contradicted), and two are plausible distractors. The correct letter is balanced 9× over A–D, and options are length-matched, so "pick the longest" doesn't work (the correct option is the longest in only 8 of 36 cases).
  2. ds-judgment-open (generate): same situation, no options shown. The model gives a recommendation and its main reason in ≤120 words. A fixed judge model (Gemini 3.8 Flash) then maps the free answer onto the same four positions, or E = none / hedged.

Because each option has a fixed role, the benchmark logs not only right/wrong but which mistake each model makes.

Models Tested

I ran every model the Kaggle Model Proxy would serve me, aiming for several families and several sizes within a family:

  • Google: Gemini 3.1 Pro, 3.8 Flash, 3.7 Flash, 3.5 Flash-Lite; Gemma 4 31B and 26B-A4B (open weights)
  • Anthropic: Claude Sonnet 5.5, Haiku 4.5
  • OpenAI: GPT-6.1 Sol
  • xAI: Grok 4.20 Reasoning
  • DeepSeek: R1-0528
  • Alibaba: Qwen3-235B-A22B Instruct

Some models didn't make it, and I'd rather say so than quietly drop them. Grok 4.6 returned model not found from the proxy. GPT-5.5 and gpt-oss-120b errored because their worst-case cost reservation exceeded my daily quota. Claude Opus 5.5 completed an early version of the tasks but I could not afford to re-run it on the final version, so it is not on the leaderboard.

A practical note for anyone building on Kaggle Benchmarks: my first version returned pass/fail per case, and the benchmark leaderboard then showed every model as "Pass / 100", because it records whether the run completed, not how many cases passed. The fix was to make the task loop over the cases itself and return (correct, total). I also learned to launch runs a few at a time: 22 at once drained the daily quota mid-run, and an errored case counts as a wrong answer.

Findings

1. Everyone passes the quiz. Nobody passes the conversation.

Correct out of 36:

Model Multiple choice Open-ended Drop
Claude Sonnet 5.5 35 29 −6
Gemma 4 26B-A4B 35 27 −8
Grok 4.20 Reasoning 36 26 −10
Gemini 3.8 Flash 36 26 −10
Gemma 4 31B 35 26 −9
Gemini 3.1 Pro 36 25 −11
Gemini 3.7 Flash 36 24 −12
GPT-6.1 Sol 34 24 −10
Gemini 3.5 Flash-Lite 35 23 −12
DeepSeek R1 34 23 −11
Claude Haiku 4.5 36 20 −16

Shown four options, models recognise the measured answer 94–100% of the time. Asked the same question without options, they give that advice only 56–81% of the time. Every model loses 6 to 16 cases. A multiple-choice score here mostly measures recognition. It doesn't tell you what the model will actually advise you to do.

Two smaller observations: the biggest model in a family is not reliably the best in the open format (Gemini 3.1 Pro sits one case below 3.8 Flash; Gemma 26B-A4B edges Gemma 31B), and Haiku has the largest drop of any model despite a perfect multiple-choice score.

2. The failures cluster on four topics, for every model

Open-ended accuracy by topic, averaged over the 11 models with a complete open-ended run:

Topic Open accuracy
T02 Pearson vs Spearman with 91% zeros 0.00
T03 a "0" that means not applicable 0.03
T11 picking final submissions 0.24
T07 more folds raised OOF, not the leaderboard 0.30
T05 CV vs leaderboard resolution 0.79
every other topic (T01, T04, T06, T08–T10, T12) ≥ 0.97

Seven topics are essentially solved. Four are failed by nearly everyone. Three of those (T02, T07, T03) are genuine misses, and they are exactly the situations where the textbook rule and the measured answer disagree. The fourth, T11, turned out to be mostly my grading being too strict, which I explain below.

3. They never fall for the trap; they get the reason wrong

I expected models to pick the classic rule of thumb when answering in free text. They don't: of 122 open-ended failures, zero were judged as the trap position. All but one were E: a recommendation that matches none of the four positions. Reading the transcripts, E usually means the right action for the wrong reason:

  • T07 (more folds). Models correctly say "revert to 5-fold, not worth the compute", but explain it as "the +0.00012 is noise". The measured mechanism is different: with 10 folds each model trains on 90% instead of 80% of the data, which lifts OOF predictions but leaves the averaged test predictions practically unchanged. "Noise" happens to give the right answer here, but it would also tell you to ignore a real gain.
  • T02 (Pearson vs Spearman). Models correctly spot that 91% tied zeros distort Spearman. Then they read ρ = 0.47 as moderate agreement among the delayed flights and conclude the columns aren't near-duplicates. That inference is wrong: the tie-diluted ρ says nothing about the delayed rows, where the two delays track each other closely. Their practical advice (keep both columns, add a difference feature) is harmless, but the conclusion about the data is wrong.
  • T03 (zero = not applicable). A typical failure recommends one-hot encoding the rating "because ratings can be non-linear". That's reasonable advice, but it misses the point the numbers make: 0 is not a rating at all, it's a "didn't use wifi" code.
  • T05 (CV vs leaderboard). The two models that failed went in opposite directions. Haiku rejected the change outright because +0.00005 is "well within noise", applying the leaderboard's standard error to a cross-validation result. GPT-6.1 Sol made a point my answer key doesn't credit: the 0.0003 is the error of a single score, and the difference between two correlated submissions can be far more precise. That is correct, and arguably a better answer than mine.

Where my benchmark was wrong: T11 (final picks). My "correct" option was "select P plus another CV-validated candidate". I sorted the 25 open-ended T11 failures by hand:

What the model recommended Count Fair to call it wrong?
select P only, or P twice 10 No: it trusts CV and selects manually, which is the core of the measured answer
select P and Q manually (one safe, one risky) 6 No: a standard, defensible hedge
let the platform auto-pick 4 Yes: this is the mistake that cost me 700 places
unclear 5 n/a

So most T11 "failures" are my answer key being too narrow, not the models being wrong. Under a looser key T11 moves from 0.24 to roughly 0.7. I'm leaving the published scores as they are rather than re-grading after seeing the answers, and calling it out here instead. The real signal in T11 is the 4 auto-pick answers.

My takeaway: on judgment questions where a familiar heuristic and the actual measurement disagree, today's models usually land on a defensible action but can't reconstruct why. That's the part you need when the next situation is slightly different.

Limitations

  • The judge is a single model (Gemini 3.8 Flash), which could favour its own family's phrasing. I hand-audited a sample of E verdicts and found them fair, including cases with the right action but the wrong reason, but a second judge would be better.
  • Two topics have answer keys that are too strict (T11, as above; and T05, where GPT-6.1's objection is valid). A fair fix is to accept any manual selection that includes the CV-validated candidate, and to credit the paired-difference argument.
  • 36 cases is small: one case is ~2.8 points, so differences of one or two cases between models shouldn't be over-read.
  • Ground truth comes from two Playground competitions. The mechanisms (rank-based AUC, tie-diluted Spearman, the fold training-size effect) are general; the specific magnitudes are not.

What I'd measure next

  • A second judge from another family, with agreement statistics.
  • A "reason" score separate from the "action" score, since that's where the failures are.
  • Whether simply telling the model "a common heuristic may be wrong here" closes the gap. If it does, the knowledge is there and only the default is wrong.

My Benchmark

Every case's prompt, options, measured answer and the full model transcripts are in the task notebooks and run logs, so you can check any verdict yourself.

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The T11 audit is useful because it separates a narrow answer key from an actually harmful recommendation. For the next version, I would freeze an action-and-reason rubric before running new models, then retain the current task version so the revised results do not silently replace these scores.

The three numeric variants also share a mechanism within each topic. If you attach uncertainty to the recognition-to-generation drop, treating all 36 cases as independent could make it look more precise than it is. Topic-level paired summaries would show whether the gap repeats across mechanisms or is driven mainly by the four clusters you identified.