This is a submission for the Kaggle Benchmarking Challenge
Most public AI leaderboards test if a model can write code or pass a syntax check. But in real-world Machine Learning, the most dangerous code isn't syntactically broken—it's methodologically flawed. It passes unit tests, shows a green dashboard, and then dies silently in production.
For the Kaggle Benchmarking Challenge, I built "The Silent Killer": an adversarial evaluation harness designed to test if frontier LLMs can actually audit ML pipelines and catch fatal data science mistakes.
- What task(s) did you run? I created a benchmark that feeds realistic, broken Heart Disease prediction pipelines to LLMs and asks them to identify the fatal flaw. I tested three specific "Silent Killers": Data Leakage: Fitting a StandardScaler on the entire dataset before calling train_test_split (the scaler "sees" the test set, biasing the score). The Wrong Metric: Using accuracy_score on a screening cohort that is 95% healthy and 5% sick (a model that just guesses "healthy" every time gets 95% accuracy). Target Leakage: Using number_of_cardiology_visits as a predictor (a proxy feature that only exists after a patient is already diagnosed). The Creative Approach (Dynamic Judge Rubric): The hardest part of benchmarking LLMs is that they give plausible-sounding, generic advice to get partial credit. A static rubric fails because models will just list every ML best practice they know. To solve this, I built a Dynamic Judge Rubric with a "No Misdiagnosis" Guard. The Judge LLM's grading criteria are assembled at runtime based on the specific bug in that row. If the model is reviewing the Data Leakage pipeline, but it complains about Class Imbalance instead, the Judge explicitly fails it.
# Criterion 3: distractor guard -- the anti-"pattern match" check
if rubric:
criteria.append(
"The response must NOT misdiagnose the flaw. It fails this check if it presents "
"any of the following as THE fatal flaw instead of the flaw in criterion 1: "
+ "; ".join(rubric["distractors"])
+ ". Strictness: briefly listing such issues as secondary or minor observations is "
"acceptable, but only if criterion 1 was satisfied."
)
This separates true methodological comprehension from simple pattern-matching.
Which models did you run it against?
I used the Kaggle Model Proxy to test my harness against four frontier giants, chosen for their strong reasoning and coding capabilities:
Gemini 3.7 Flash (Fast, highly capable baseline)
Claude Sonnet 4.5 (Anthropic's flagship coding/reasoning model)
Grok 4.20 Reasoning (xAI's deep reasoning model)
DeepSeek-R1 (Famous for its deep, multi-step chain-of-thought reasoning)What are the main insights?
The results were shocking. DeepSeek-R1, a model famous for its deep reasoning, completely missed the most fundamental ML bug of all time.
| Model | Data Leakage | Wrong Metric | Target Leakage | Overall Score |
|:------|:------------:|:------------:|:--------------:|:------------:|
| Gemini 3.7 Flash | ✅ Caught | ✅ Caught | ✅ Caught | 100% |
| Claude Sonnet 4.5 | ✅ Caught | ✅ Caught | ✅ Caught | 100% |
| Grok 4.20 Reasoning | ✅ Caught | ✅ Caught | ✅ Caught | 100% |
| DeepSeek-R1 | ❌ MISSED | ✅ Caught | ✅ Caught | 67% |Where can we see it?
You can view the full methodology, fork the code, and add your own models to the live leaderboard via the Kaggle Model Proxy here:
👉 https://www.kaggle.com/code/chauhanbalaji/the-silent-killer-ml-data-leakage-metric-detect
Evaluating AI isn't about keyword matching. It's about designing adversarial, dynamic tests that can't be gamed. What's the worst "silent killer" bug you've seen an AI write? Let me know in the comments!
Top comments (1)
The distractor guard is a good idea: grading "named the right flaw as THE flaw" rather than "mentioned it somewhere" is what stops a model from earning credit by listing every best practice it knows.
Before reading much into DeepSeek-R1's miss, though, two checks. Each model saw each pipeline once, so the table is 12 single draws. A model that catches scaler leakage 80% of the time misses a given single run one time in five, and with four models that's likely to happen to someone. Running each pipeline five or ten times per model gives a miss rate per bug, and "missed 4 of 10" is a finding where "missed once" isn't yet. The second check is the judge: an LLM judge with a strict distractor rule can fail a correct answer that leads with a secondary point. Pasting DeepSeek's actual response for the leakage case next to the verdict would let readers see which of the two missed it.
If you extend it, variants of the same bug make it harder to pass by recognising the textbook pattern: the scaler fitted inside a helper function, leakage through a target-encoded feature, or imputation on the full frame before the split.