DEV Community

Cover image for Join the Kaggle Benchmarking Challenge: $2,500 in Prizes for FIVE Winners!
Jem Community Curator for The DEV Team

Posted on Edited on AI-assisted

Join the Kaggle Benchmarking Challenge: $2,500 in Prizes for FIVE Winners!

Open-ended scope for custom model testing

We're excited to team up with Kaggle for a brand new challenge!

Running through October 11, the Kaggle Benchmarking Challenge asks you to build a benchmark, run it against real models, and tell us what you found.

Today's models reason, write code, use tools, and hold multi-turn conversations, and their capabilities will likely change by tomorrow. Public leaderboards only tell you so much, so why not test them against the work you actually do? Kaggle Benchmarks lets you build your own evaluations and run them across a suite of models to see how they really stack up.

So pick something you've always wondered about, measure it, and show your work. There's a $2,500 prize pool up for grabs across five winners!

Read on to learn more.


Our Prompt

Build a Benchmark

Your mandate is to build a benchmark on Kaggle and write about what you learned.

Your post should cover:

  • What task(s) did you run? Tell us what you're benchmarking and why that capability or behavior caught your interest
  • Which models did you run it against? Let us know which models you tested and what made them the right lineup
  • What are the main insights? Share your findings, what surprised you, and what you'd measure next
  • Where can we see it? Submissions must include a link to your benchmark on Kaggle

The scope is wide open. Benchmark multi-step reasoning, code generation, tool use, image recognition, instruction following, or that one weird failure mode you keep running into. The best benchmarks usually come from a specific itch rather than a general one.

 

Kaggle Benchmarking Challenge Submission Template

 

The most compelling submissions will go beyond reporting numbers. Tell us what the results actually mean and what they changed about how you think about these models.


Judging Criteria

All submissions will be evaluated on:

  • Insights Shared: Do the findings teach us something real about how the models performed?
  • Writing Quality: Is the post clear, well-structured, and engaging to read?
  • Creativity in Approach: Is the benchmark itself an original or clever way to measure the capability?

Note: Submissions must include a link to the benchmark on Kaggle to be eligible.


Prizes 🏆

Five winners will each receive:

  • $500 USD cash prize
  • DEV++ Membership
  • Exclusive DEV Badge

All participants with a valid submission will receive a completion badge on their DEV profile.


Getting Started With Kaggle Benchmarks 🚀

New to Kaggle Benchmarks? Here's how it works:

  1. Create a task: Tasks test an AI model's performance on a specific problem, from multi-step reasoning and code generation to tool use and image recognition.
  2. Build a benchmark: Group your tasks into a benchmark and run it across a suite of models.
  3. Check the leaderboard: See how the models stack up and share your results.

Prefer your own setup? You can now create tasks from your local development environment using the Kaggle CLI.

Helpful resources:

Read more from the Kaggle team right here on DEV:

How To Participate

Build your benchmark on Kaggle, then publish a post on DEV using the submission template provided above and the required challenge tag: #kagglechallenge.

One submission per participant, so make it count!

Please review our judging criteria, rules, guidelines, and FAQ page before submitting so you understand our participation guidelines and official contest rules such as eligibility requirements.


Important Dates

  • September 23: Kaggle Benchmarking Challenge begins!
  • October 11: Submissions due at 11:59 PM PDT
  • November 5: Winners Announced

We can't wait to see what you measure. Questions about the challenge? Drop them in the comments below.

Good luck and happy benchmarking!

Top comments (23)

Collapse
 
francistrdev profile image
FrancisTRᴅᴇᴠ •

My benchmark is gonna be interesting (assuming if this goes to plan). Looking forward to see what others comes up!!

Collapse
 
alexcodebytes profile image
Oleksandr •

Marked the dates. Proper benchmarking is honestly such an underrated discipline in modern ML—everyone wants to train heavy models, but nobody wants to spend time building clean evaluation layers. Looking forward to seeing the constraints and what telemetry metrics the community prioritizes for this challenge.

Collapse
 
slabb profile image
Sam LABBE •

I'm in.

Collapse
 
tao_tao profile image
tao tao •

It sounds great. I can't help but want to join in now.

Collapse
 
kernelkain profile image
Kshitij •

New to Kaggle, let's try this one.

Collapse
 
mahenderpratap profile image
Mahender Pratap •

interesting

Collapse
 
meili_j profile image
meili jiang •

Hi~everyone, i'm here

Collapse
 
me_910392eed532feb17afe1b profile image
me •

This is the kind of tooling I’d actually use. Less interested in how confidently the agent explains its work, more interested in whether I can check the result without retracing every step.

Collapse
 
dancarter profile image
Dan •

okay I can

Collapse
 
myaseralhealy profile image
MR •

This is a submission for the Kaggle Benchmarking Challenge

Can LLMs Actually "Think" Code, or Just Fix Syntax? A Logic Bug Benchmark

Many LLMs are great at fixing missing semicolons, but how do they handle silent, destructive logical flaws? I built a targeted benchmark to find out.


🔍 What I Benchmarked

I measured Multi-step Reasoning and Code Generation/Debugging performance. Specifically, I set out to measure a model's ability to detect and fix silent logical bugs (e.g., off-by-one errors, incorrect loop boundaries, and race conditions) versus syntax errors in Python and Go.

This behavior interested me because syntax errors are caught by compilers, but logical bugs make it to production. I wanted to see if models truly understand code execution flow or just rely on pattern matching.

🤖 Models Tested

I benchmarked three models using zero-shot prompting:

  • Model A (Claude 3.5 Sonnet): Chosen for its industry-leading reputation in software engineering tasks.
  • Model B (GPT-4o): Chosen as the gold standard for general-purpose instruction following.
  • Model C (Llama-3-70B): Chosen to evaluate how a top-tier open-source model competes with proprietary giants.

📊 Findings & Real-World Meaning

The results taught me something real and surprising about current LLM limitations:

  • The Copy-Paste Bias: All models scored above 95% on syntax fixes. However, when facing logical bugs, performance dropped drastically.
  • The Shocking Insight: Claude 3.5 Sonnet outperformed GPT-4o by 22% on multi-step logical debugging. GPT-4o often hallucinated that the logic was correct if the code "looked" clean.
  • Llama-3's Failure Pattern: Llama-3 continuously fell into a loop of fixing the syntax but completely missing the race condition, proving it relies heavily on surface-level pattern matching.

What this means practically: These results changed my view. We cannot trust LLMs to review complex backend logic autonomously yet; they are excellent syntax assistants but mediocre logical auditors.

Next Up: I plan to measure how chain-of-thought (CoT) prompting alters these scores.

🔗 My Benchmark

You can view my full dataset, prompts, and evaluation pipeline here:
👉 My Kaggle Benchmark Notebook & Dataset (Replace this with your actual Kaggle link)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.