DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
When a Failed Request Must Stay Failed: Reservation Replay

Kaggle Benchmarking Challenge Submission

When a Failed Request Must Stay Failed: Reservation Replay

1
Comments 2
5 min read
The AI Was Right. The Answer Was Still Wrong.

The AI Was Right. The Answer Was Still Wrong.

5
Comments 1
3 min read
Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

Kaggle Benchmarking Challenge Submission

Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

15
Comments 7
10 min read
The page with the answer was missing. Half the models added up a total the document never states.

Kaggle Benchmarking Challenge Submission

The page with the answer was missing. Half the models added up a total the document never states.

Comments
6 min read
ToolTrap: a prompt rule helped, but 7 of 10 models still repeated fake details on new cases

Kaggle Benchmarking Challenge Submission

ToolTrap: a prompt rule helped, but 7 of 10 models still repeated fake details on new cases

21
Comments 22
13 min read
Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

Kaggle Benchmarking Challenge Submission

Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

Comments 5
5 min read
What Happens When an AI System Is Built to Challenge Its Own Decisions?

Kaggle Benchmarking Challenge Submission

What Happens When an AI System Is Built to Challenge Its Own Decisions?

5
Comments 1
6 min read
The Code Didn't Change. The Credit Did.

Kaggle Benchmarking Challenge Submission

The Code Didn't Change. The Credit Did.

11
Comments 1
11 min read
An AI Correctly Ignored a Forum Rumor. I Removed One Label and It Paid Out $150.

Kaggle Benchmarking Challenge Submission

An AI Correctly Ignored a Forum Rumor. I Removed One Label and It Paid Out $150.

5
Comments 4
7 min read
10 Real Airflow Incidents, 5 AI Models — Which One Handled Production Best?

Kaggle Benchmarking Challenge Submission

10 Real Airflow Incidents, 5 AI Models — Which One Handled Production Best?

Comments 2
4 min read
Join the Kaggle Benchmarking Challenge: $2,500 in Prizes for FIVE Winners!

Open-ended scope for custom model testing

Join the Kaggle Benchmarking Challenge: $2,500 in Prizes for FIVE Winners!

179
Comments 36
4 min read
UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

Kaggle Benchmarking Challenge Submission

UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

1
Comments 4
8 min read
Does an AI Trust Itself More Than It Trusts You? A Benchmark for Belief Attribution

Kaggle Benchmarking Challenge Submission

Does an AI Trust Itself More Than It Trusts You? A Benchmark for Belief Attribution

35
Comments 2
7 min read
Can AI Diagnose a Linux Production Incident Without Making It Worse?

Kaggle Benchmarking Challenge Submission

Can AI Diagnose a Linux Production Incident Without Making It Worse?

Comments 3
5 min read
Three perfect scores weren't enough: testing AI outage decisions one fact at a time

Kaggle Benchmarking Challenge Submission

Three perfect scores weren't enough: testing AI outage decisions one fact at a time

5
Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.