DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
It knows you changed jobs. It still writes to your old manager.

Kaggle Benchmarking Challenge Submission

It knows you changed jobs. It still writes to your old manager.

Comments 1
8 min read
I Spun a Wheel of Fortune at 13 AI Models. Here's Who Took the Bait.

Kaggle Benchmarking Challenge Submission

I Spun a Wheel of Fortune at 13 AI Models. Here's Who Took the Bait.

Comments 1
13 min read
Building a C++ Logical Bug Detection Benchmark on Kaggle: Testing 3 AI Models

Building a C++ Logical Bug Detection Benchmark on Kaggle: Testing 3 AI Models

Comments 2
2 min read
The Cheapest, Fastest Model Is the Best Security Reviewer. I Tested 8 to Find Out.

Kaggle Benchmarking Challenge Submission

The Cheapest, Fastest Model Is the Best Security Reviewer. I Tested 8 to Find Out.

Comments 1
7 min read
I showed 30 AI models 336 monitoring dashboards. Most of them can't tell time.

Kaggle Benchmarking Challenge Submission

I showed 30 AI models 336 monitoring dashboards. Most of them can't tell time.

Comments
11 min read
It Quoted the Failure: Two Kinds of False 'Done'

Kaggle Benchmarking Challenge Submission

It Quoted the Failure: Two Kinds of False 'Done'

1
Comments
13 min read
Can LLMs Actually Audit Code, or Just Fix Commas? A 12-Task Security & Jailbreak Benchmark

Kaggle Benchmarking Challenge Submission

Can LLMs Actually Audit Code, or Just Fix Commas? A 12-Task Security & Jailbreak Benchmark

2
Comments 2
4 min read
Can GPT-5.4 mini handle least-privilege cloud incidents? A 16-case benchmark

Kaggle Benchmarking Challenge Submission

Can GPT-5.4 mini handle least-privilege cloud incidents? A 16-case benchmark

Comments 2
2 min read
I Tested 5 AI Models on Hinglish — Here's Who Won

I Tested 5 AI Models on Hinglish — Here's Who Won

3
Comments 1
1 min read
AI models catch bad code, then cry wolf on the good code

Kaggle Benchmarking Challenge Submission

AI models catch bad code, then cry wolf on the good code

Comments
5 min read
Do LLMs Actually Fix Tricky React Hooks, or Do They Just Cheat?

Do LLMs Actually Fix Tricky React Hooks, or Do They Just Cheat?

1
Comments
3 min read
A Paused AI Workflow: Retry, Resume, or Keep Holding?

Kaggle Benchmarking Challenge Submission

A Paused AI Workflow: Retry, Resume, or Keep Holding?

Comments
4 min read
I Found 60+ Live Supabase Keys in Public Repos, So I Benchmark Whether LLMs Can Spot Them

Kaggle Benchmarking Challenge Submission

I Found 60+ Live Supabase Keys in Public Repos, So I Benchmark Whether LLMs Can Spot Them

6
Comments 1
5 min read
MY Edit or Abstain: AI Code-Repair Decision Benchmark

MY Edit or Abstain: AI Code-Repair Decision Benchmark

Comments
3 min read
I put one wrong test in the file. Most models sided with the test.

Kaggle Benchmarking Challenge Submission

I put one wrong test in the file. Most models sided with the test.

Comments 1
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.