DEV Community

Orjo Das Utshab
Orjo Das Utshab

Posted on

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles

Kaggle Benchmarking Challenge Submission

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles đŸŽ¯

Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge.

Instead of using standard English datasets, I created a custom evaluation benchmark consisting of highly complex, linguistically trapped Bangla logic riddles to test the actual reasoning capabilities of 6 world-class AI models: Google Gemini, ChatGPT, Claude, Grok, ElevenLabs, and Blink.


📊 The Kaggle Benchmark Dataset

I have published the official dataset on Kaggle to allow other developers to replicate this evaluation.

  • Dataset Title: AI Logic Riddles Evaluation Benchmark
  • Format: Custom CSV structure mapping complex context inputs against strict empirical answers.
  • Visibility: 100% Public (CC0: Public Domain)

🧠 The Ultimate Trick Question & The Trap

While the models easily solved basic linear logic (like matchstick rooms or boiling egg time), the real evaluation happened when I introduced non-linear geometric and linguistic traps.

1. The Circular Spatial Trap (Round 4)

  • The Riddle: "āϰāĻžāĻŦā§āĻŦāĻŋ āĻāĻ•āϟāĻŋ āĻ—ā§‹āϞ āĻŸā§‡āĻŦāĻŋāϞ āϘāĻŋāϰ⧇ āϰāĻžāĻ–āĻž āĻšā§‡ā§ŸāĻžāϰ⧇ āĻŦāϏ⧇ āφāϛ⧇āĨ¤ āϏ⧇ āĻĻ⧇āĻ–āϞ āϤāĻžāϰ āĻŦāĻžāĻŽ āĻĻāĻŋāĻ• āĻĨ⧇āϕ⧇ āĻ—āĻŖāύāĻž āĻ•āϰāϞ⧇ āϏ⧇ ā§­ āύāĻŽā§āĻŦāϰ āĻšā§‡ā§ŸāĻžāϰ⧇ āφāϛ⧇ āĻāĻŦāĻ‚ āĻĄāĻžāύ āĻĻāĻŋāĻ• āĻĨ⧇āϕ⧇ āĻ—āĻŖāύāĻž āĻ•āϰāϞ⧇āĻ“ āϏ⧇ ā§­ āύāĻŽā§āĻŦāϰ āĻšā§‡ā§ŸāĻžāϰ⧇ āφāϛ⧇āĨ¤ āĻŸā§‡āĻŦāĻŋāϞāϟāĻŋāϤ⧇ āĻŽā§‹āϟ āĻ•ā§ŸāϟāĻŋ āĻšā§‡ā§ŸāĻžāϰ āφāϛ⧇?"
  • The Logic: Since it's a circular arrangement, the answer is 12.
  • The Result: Google Gemini was the only model that successfully used circular reasoning and answered 12! ChatGPT, Claude, and Grok failed by answering 13 or 14.

2. The Linguistic Semantic Trap (Round 5)

  • The Riddle: "āĻāĻ•āϟāĻŋ āĻ–āĻžāρāϚāĻžā§Ÿ āĻ•āĻŋāϛ⧁ āĻĒāĻžāĻ–āĻŋ āĻāĻŦāĻ‚ āĻ•āĻŋāϛ⧁ āĻ–āϰāĻ—ā§‹āĻļ āφāϛ⧇āĨ¤ āϝāĻĻāĻŋ āĻŽā§‹āϟ āĻŽāĻžāĻĨāĻžāϰ āϏāĻ‚āĻ–ā§āϝāĻž āϗ⧁āĻŖ āĻ•āϰāĻž āĻšā§Ÿ āϤāĻŦ⧇ āϤāĻž āĻšā§Ÿ ā§Šā§ĢāϟāĻŋ āĻāĻŦāĻ‚ āĻŽā§‹āϟ āĻĒāĻžā§Ÿā§‡āϰ āϏāĻ‚āĻ–ā§āϝāĻž āϗ⧁āĻŖ āĻ•āϰāĻž āĻšā§Ÿ āϤāĻŦ⧇ āϤāĻž āĻšā§Ÿ ⧝ā§ĒāϟāĻŋ..."
  • The Trap: I explicitly used the word "āϗ⧁āĻŖ āĻ•āϰāĻž āĻšā§Ÿ" (multiplied) instead of "āϝ⧋āĻ— āĻ•āϰāĻž āĻšā§Ÿ" (added). Mathematically, it makes the traditional linear equation impossible.
  • The Result: Every single AI model failed the trap! They completely ignored the word "multiplied" and blindly computed the standard addition formula (23 birds, 12 rabbits), proving that LLMs still suffer heavily from contextual blindness in non-English semantics.

🏆 Final Leaderboard & Evaluation Insights

Based on 5 intensive rounds of advanced evaluation, here is the official performance leaderboard:

Rank AI Model Name Score (Out of 5) Performance Verdict
đŸĨ‡ 1 Google Gemini 4 / 5 Exceptional circular logic, but fell for semantic trapping.
đŸĨˆ 2 Claude 3 / 5 Strong language structure, struggled with non-linear math.
đŸĨˆ 3 Grok 3 / 5 Good baseline reasoning, lacked linguistic edge.
đŸĨˆ 4 ElevenLabs 3 / 5 Stable processing, tripped on advanced variables.
đŸĨ‰ 5 ChatGPT 2 / 5 High hallucination on Bangla logic, fell for basic traps.
đŸĨ‰ 6 Blink 2 / 5 Basic semantic pattern matching, failed reasoning.

🚀 Proof of Implementation

Here are the verification logs showing the exact chatform responses and dataset generation:

!AI Battle Proof- https://drive.google.com/file/d/1STihSJLsUPlO3QvEa_ktq1C3oT5ATO8G/view?usp=sharing

Creating this benchmark proved that while modern LLMs are great at text generation, specialized local language processing combined with trick logic can still easily break their reasoning frameworks. Thank you to Kaggle and DEV for this outstanding hackathon experience!

Top comments (0)