DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Hook constraint benchmark kaggle-challenge DONE !!!

Kaggle Benchmarking Challenge Submission

Hook constraint benchmark kaggle-challenge DONE !!!

Comments
3 min read
Small AI Models Can See the Trap. They Fall In Anyway

Kaggle Benchmarking Challenge Submission

Small AI Models Can See the Trap. They Fall In Anyway

1
Comments
5 min read
Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Kaggle Benchmarking Challenge Submission

Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Comments
4 min read
Do LLMs Actually Check Their Tools? I Built a Benchmark That Lies to Them

Kaggle Benchmarking Challenge Submission

Do LLMs Actually Check Their Tools? I Built a Benchmark That Lies to Them

2
Comments 5
8 min read
100% vuln detection wasn't enough: measuring whether AI respects the patch

Kaggle Benchmarking Challenge Submission

100% vuln detection wasn't enough: measuring whether AI respects the patch

7
Comments 4
6 min read
LAW-N Real-World Data Layer

LAW-N Real-World Data Layer

3
Comments
8 min read
Don't Take Orders From the Internet: Benchmarking 5 LLMs Against Indirect Prompt Injection

Kaggle Benchmarking Challenge Submission

Don't Take Orders From the Internet: Benchmarking 5 LLMs Against Indirect Prompt Injection

Comments
7 min read
A CSV benchmark with two zero scores and a useful diagnostic

Kaggle Benchmarking Challenge Submission

A CSV benchmark with two zero scores and a useful diagnostic

Comments 1
7 min read
I Surveyed 123 People in India to Benchmark Frontier AI

Kaggle Benchmarking Challenge Submission

I Surveyed 123 People in India to Benchmark Frontier AI

26
Comments
5 min read
DarijaBench: Do AI Models Actually Understand Moroccan Darija?

Kaggle Benchmarking Challenge Submission

DarijaBench: Do AI Models Actually Understand Moroccan Darija?

1
Comments 2
2 min read
Smaller models often read URLs like Python, not like fetch(). I benchmarked where the API key leaks

kagglechallenge

Kaggle Benchmarking Challenge Submission

Smaller models often read URLs like Python, not like fetch(). I benchmarked where the API key leaks

8
Comments 2
10 min read
Day 1: Most of My Bugs Looked Like Model Behaviour

Day 1: Most of My Bugs Looked Like Model Behaviour

Comments
5 min read
The explanation was right. The policy ID was wrong.

Kaggle Benchmarking Challenge Submission

The explanation was right. The policy ID was wrong.

Comments
5 min read
1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.

Kaggle Benchmarking Challenge Submission

1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.

1
Comments 1
5 min read
Promise Is Not Payment: Verification Errors That Amount Accuracy Misses

Kaggle Benchmarking Challenge Submission

Promise Is Not Payment: Verification Errors That Amount Accuracy Misses

Comments 2
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.