DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Only What I Asked

Only What I Asked

Comments
3 min read
Hook constraint benchmark kaggle-challenge DONE !!!

Kaggle Benchmarking Challenge Submission

Hook constraint benchmark kaggle-challenge DONE !!!

Comments
3 min read
The AI Was Right. The Answer Was Still Wrong.

The AI Was Right. The Answer Was Still Wrong.

5
Comments 1
3 min read
UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

Kaggle Benchmarking Challenge Submission

UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

Comments
8 min read
Don't Take Orders From the Internet: Benchmarking 5 LLMs Against Indirect Prompt Injection

Kaggle Benchmarking Challenge Submission

Don't Take Orders From the Internet: Benchmarking 5 LLMs Against Indirect Prompt Injection

Comments
5 min read
Does an AI Trust Itself More Than It Trusts You? A Benchmark for Belief Attribution

Kaggle Benchmarking Challenge Submission

Does an AI Trust Itself More Than It Trusts You? A Benchmark for Belief Attribution

20
Comments 1
7 min read
# **Titanic – Machine Learning From Disaster: A Complete Project Overview**

Kaggle Benchmarking Challenge Submission

# **Titanic – Machine Learning From Disaster: A Complete Project Overview**

Comments
4 min read
Small AI Models Can See the Trap. They Fall In Anyway

Kaggle Benchmarking Challenge Submission

Small AI Models Can See the Trap. They Fall In Anyway

1
Comments
5 min read
Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Kaggle Benchmarking Challenge Submission

Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Comments
4 min read
100% vuln detection wasn't enough: measuring whether AI respects the patch

Kaggle Benchmarking Challenge Submission

100% vuln detection wasn't enough: measuring whether AI respects the patch

7
Comments 4
6 min read
Can AI Diagnose a Linux Production Incident Without Making It Worse?

Kaggle Benchmarking Challenge Submission

Can AI Diagnose a Linux Production Incident Without Making It Worse?

Comments 2
5 min read
Three perfect scores weren't enough: testing AI outage decisions one fact at a time

Kaggle Benchmarking Challenge Submission

Three perfect scores weren't enough: testing AI outage decisions one fact at a time

5
Comments
5 min read
Breaking Character: How "Thinking" AI Survives Villain Roleplay Traps

Breaking Character: How "Thinking" AI Survives Villain Roleplay Traps

Comments
2 min read
How AI Models "See" Hidden Meaning: A Beginner's Subtext Benchmark

How AI Models "See" Hidden Meaning: A Beginner's Subtext Benchmark

2
Comments
1 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.