DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
I put one wrong test in the file. Most models sided with the test.

Kaggle Benchmarking Challenge Submission

I put one wrong test in the file. Most models sided with the test.

Comments 1
6 min read
Only What I Asked

Only What I Asked

Comments
3 min read
Hook constraint benchmark kaggle-challenge DONE !!!

Kaggle Benchmarking Challenge Submission

Hook constraint benchmark kaggle-challenge DONE !!!

Comments
3 min read
UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

Kaggle Benchmarking Challenge Submission

UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

Comments
8 min read
10 Real Airflow Incidents, 5 AI Models — Which One Handled Production Best?

Kaggle Benchmarking Challenge Submission

10 Real Airflow Incidents, 5 AI Models — Which One Handled Production Best?

Comments 1
4 min read
Ask ten LLMs for a Blender 5.0 script. Then run it in Blender 5.0.

Kaggle Benchmarking Challenge Submission

Ask ten LLMs for a Blender 5.0 script. Then run it in Blender 5.0.

3
Comments 1
9 min read
Small AI Models Can See the Trap. They Fall In Anyway

Kaggle Benchmarking Challenge Submission

Small AI Models Can See the Trap. They Fall In Anyway

1
Comments
5 min read
Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Kaggle Benchmarking Challenge Submission

Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Comments
4 min read
Do LLMs Actually Check Their Tools? I Built a Benchmark That Lies to Them

Kaggle Benchmarking Challenge Submission

Do LLMs Actually Check Their Tools? I Built a Benchmark That Lies to Them

2
Comments 5
8 min read
100% vuln detection wasn't enough: measuring whether AI respects the patch

Kaggle Benchmarking Challenge Submission

100% vuln detection wasn't enough: measuring whether AI respects the patch

7
Comments 4
6 min read
Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

Kaggle Benchmarking Challenge Submission

Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

Comments 5
3 min read
Can AI Diagnose a Linux Production Incident Without Making It Worse?

Kaggle Benchmarking Challenge Submission

Can AI Diagnose a Linux Production Incident Without Making It Worse?

Comments 2
5 min read
Chain-of-Thought Faithfulness: Toggling 'Reasoning Mode' Made One Model 5x More Likely to Follow Its Own Mistakes

Kaggle Benchmarking Challenge Submission

Chain-of-Thought Faithfulness: Toggling 'Reasoning Mode' Made One Model 5x More Likely to Follow Its Own Mistakes

44
Comments 20
5 min read
LAW-N Real-World Data Layer

LAW-N Real-World Data Layer

3
Comments
8 min read
Don't Take Orders From the Internet: Benchmarking 5 LLMs Against Indirect Prompt Injection

Kaggle Benchmarking Challenge Submission

Don't Take Orders From the Internet: Benchmarking 5 LLMs Against Indirect Prompt Injection

Comments
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.