DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
# **Titanic – Machine Learning From Disaster: A Complete Project Overview**

Kaggle Benchmarking Challenge Submission

# **Titanic – Machine Learning From Disaster: A Complete Project Overview**

1
Comments
4 min read
Chain-of-Thought Faithfulness: Toggling 'Reasoning Mode' Made One Model 5x More Likely to Follow Its Own Mistakes

Kaggle Benchmarking Challenge Submission

Chain-of-Thought Faithfulness: Toggling 'Reasoning Mode' Made One Model 5x More Likely to Follow Its Own Mistakes

58
Comments 28
5 min read
Breaking Character: How "Thinking" AI Survives Villain Roleplay Traps

Breaking Character: How "Thinking" AI Survives Villain Roleplay Traps

Comments
2 min read
How AI Models "See" Hidden Meaning: A Beginner's Subtext Benchmark

How AI Models "See" Hidden Meaning: A Beginner's Subtext Benchmark

2
Comments
1 min read
I Benchmarked 10 AI Models to Reconstruct Real Cyber Attacks

Kaggle Benchmarking Challenge Submission

I Benchmarked 10 AI Models to Reconstruct Real Cyber Attacks

6
Comments 3
13 min read
Can LLMs audit multi-agent prompts? I made them grade my robot football team.

Kaggle Benchmarking Challenge Submission

Can LLMs audit multi-agent prompts? I made them grade my robot football team.

Comments
5 min read
8 LLMs, 480 Questions, 1 Kaggle Benchmark: Who Can Explain a Traffic Drop?

Kaggle Benchmarking Challenge Submission

8 LLMs, 480 Questions, 1 Kaggle Benchmark: Who Can Explain a Traffic Drop?

9
Comments 13
11 min read
Does your model know when it doesn't know? A benchmark for the ESCALATE answer

Kaggle Benchmarking Challenge Submission

Does your model know when it doesn't know? A benchmark for the ESCALATE answer

1
Comments 4
2 min read
Same Patient, Conflicting Documents: Can AI Preserve the Evidence?

Kaggle Benchmarking Challenge Submission

Same Patient, Conflicting Documents: Can AI Preserve the Evidence?

7
Comments
13 min read
The Plausible PR: I Gave 4 LLMs 10 Sneaky Refactors, and They All Missed the Same Bug

Kaggle Benchmarking Challenge Submission

The Plausible PR: I Gave 4 LLMs 10 Sneaky Refactors, and They All Missed the Same Bug

2
Comments 2
3 min read
Ask ten LLMs for a Blender 5.0 script. Then run it in Blender 5.0.

Kaggle Benchmarking Challenge Submission

Ask ten LLMs for a Blender 5.0 script. Then run it in Blender 5.0.

6
Comments 2
9 min read
ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM

Kaggle Benchmarking Challenge Submission

ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM

10
Comments
14 min read
AI reviewers catch Italian UI bugs, then flag the good fixes too

Kaggle Benchmarking Challenge Submission

AI reviewers catch Italian UI bugs, then flag the good fixes too

Comments
5 min read
I Benchmarked Whether AI Models Forget Corrections. The Answer Surprised Me.

Kaggle Benchmarking Challenge Submission

I Benchmarked Whether AI Models Forget Corrections. The Answer Surprised Me.

Comments 1
4 min read
Playoff Probability Calibration — LLMs vs. a Real Monte Carlo Model

Playoff Probability Calibration — LLMs vs. a Real Monte Carlo Model

Comments 1
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.