Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles đ¯
Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge.
Instead of using standard English datasets, I created a custom evaluation benchmark consisting of highly complex, linguistically trapped Bangla logic riddles to test the actual reasoning capabilities of 6 world-class AI models: Google Gemini, ChatGPT, Claude, Grok, ElevenLabs, and Blink.
đ The Kaggle Benchmark Dataset
I have published the official dataset on Kaggle to allow other developers to replicate this evaluation.
-
Dataset Title:
AI Logic Riddles Evaluation Benchmark - Format: Custom CSV structure mapping complex context inputs against strict empirical answers.
- Visibility: 100% Public (CC0: Public Domain)
đ§ The Ultimate Trick Question & The Trap
While the models easily solved basic linear logic (like matchstick rooms or boiling egg time), the real evaluation happened when I introduced non-linear geometric and linguistic traps.
1. The Circular Spatial Trap (Round 4)
- The Riddle: "āϰāĻžāĻŦā§āĻŦāĻŋ āĻāĻāĻāĻŋ āĻā§āϞ āĻā§āĻŦāĻŋāϞ āĻāĻŋāϰ⧠āϰāĻžāĻāĻž āĻā§ā§āĻžāϰ⧠āĻŦāϏ⧠āĻāĻā§āĨ¤ āϏ⧠āĻĻā§āĻāϞ āϤāĻžāϰ āĻŦāĻžāĻŽ āĻĻāĻŋāĻ āĻĨā§āĻā§ āĻāĻŖāύāĻž āĻāϰāϞ⧠āϏ⧠⧠āύāĻŽā§āĻŦāϰ āĻā§ā§āĻžāϰ⧠āĻāĻā§ āĻāĻŦāĻ āĻĄāĻžāύ āĻĻāĻŋāĻ āĻĨā§āĻā§ āĻāĻŖāύāĻž āĻāϰāϞā§āĻ āϏ⧠⧠āύāĻŽā§āĻŦāϰ āĻā§ā§āĻžāϰ⧠āĻāĻā§āĨ¤ āĻā§āĻŦāĻŋāϞāĻāĻŋāϤ⧠āĻŽā§āĻ āĻā§āĻāĻŋ āĻā§ā§āĻžāϰ āĻāĻā§?"
- The Logic: Since it's a circular arrangement, the answer is 12.
- The Result: Google Gemini was the only model that successfully used circular reasoning and answered 12! ChatGPT, Claude, and Grok failed by answering 13 or 14.
2. The Linguistic Semantic Trap (Round 5)
- The Riddle: "āĻāĻāĻāĻŋ āĻāĻžāĻāĻāĻžā§ āĻāĻŋāĻā§ āĻĒāĻžāĻāĻŋ āĻāĻŦāĻ āĻāĻŋāĻā§ āĻāϰāĻā§āĻļ āĻāĻā§āĨ¤ āϝāĻĻāĻŋ āĻŽā§āĻ āĻŽāĻžāĻĨāĻžāϰ āϏāĻāĻā§āϝāĻž āĻā§āĻŖ āĻāϰāĻž āĻšā§ āϤāĻŦā§ āϤāĻž āĻšā§ ā§Šā§ĢāĻāĻŋ āĻāĻŦāĻ āĻŽā§āĻ āĻĒāĻžā§ā§āϰ āϏāĻāĻā§āϝāĻž āĻā§āĻŖ āĻāϰāĻž āĻšā§ āϤāĻŦā§ āϤāĻž āĻšā§ ā§¯ā§ĒāĻāĻŋ..."
- The Trap: I explicitly used the word "āĻā§āĻŖ āĻāϰāĻž āĻšā§" (multiplied) instead of "āϝā§āĻ āĻāϰāĻž āĻšā§" (added). Mathematically, it makes the traditional linear equation impossible.
- The Result: Every single AI model failed the trap! They completely ignored the word "multiplied" and blindly computed the standard addition formula (23 birds, 12 rabbits), proving that LLMs still suffer heavily from contextual blindness in non-English semantics.
đ Final Leaderboard & Evaluation Insights
Based on 5 intensive rounds of advanced evaluation, here is the official performance leaderboard:
| Rank | AI Model Name | Score (Out of 5) | Performance Verdict |
|---|---|---|---|
| đĨ 1 | Google Gemini | 4 / 5 | Exceptional circular logic, but fell for semantic trapping. |
| đĨ 2 | Claude | 3 / 5 | Strong language structure, struggled with non-linear math. |
| đĨ 3 | Grok | 3 / 5 | Good baseline reasoning, lacked linguistic edge. |
| đĨ 4 | ElevenLabs | 3 / 5 | Stable processing, tripped on advanced variables. |
| đĨ 5 | ChatGPT | 2 / 5 | High hallucination on Bangla logic, fell for basic traps. |
| đĨ 6 | Blink | 2 / 5 | Basic semantic pattern matching, failed reasoning. |
đ Proof of Implementation
Here are the verification logs showing the exact chatform responses and dataset generation:
!AI Battle Proof- https://drive.google.com/file/d/1STihSJLsUPlO3QvEa_ktq1C3oT5ATO8G/view?usp=sharing
Creating this benchmark proved that while modern LLMs are great at text generation, specialized local language processing combined with trick logic can still easily break their reasoning frameworks. Thank you to Kaggle and DEV for this outstanding hackathon experience!
Top comments (0)