DEV Community

Arshad Ansari
Arshad Ansari

Posted on

I Tested 5 AI Models on Hinglish — Here's Who Won

I'm a CSE student from Maharashtra, India. I code everything from my mobile phone. And I speak Hinglish — that beautiful mix of Hindi and English that 600 million Indians use every day.

So I had one question: Do AI models actually understand us?

Not "formal English" us. I mean the real us — "Yaar ye loop khatam hi nahi ho raha, kuch kar" us.

I built a benchmark on Kaggle to find out.


What is Hinglish?

Hinglish is how most Indian developers actually communicate — a natural mix of Hindi and English in the same sentence.

Examples:

  • "Yaar, mera code kuch kaam nahi kar raha. Fix kar de."
  • "Bhai ye function output nahi de raha, dekh zara."
  • "Kal se ye wala loop khatam hi nahi ho raha."

This is how we talk to each other. But do AI models understand this? That's what I tested.


The Benchmark

I created a Hinglish Code Understanding Benchmark on Kaggle with tasks across 3 categories:

🐛 Category 1: Hinglish Code Debugging

The model gets broken Python code with instructions in Hinglish. Can it understand the problem and fix the code?

Example prompt:

"Yaar, mera code kuch kaam nahi kar raha. Zara fix kar de:
def add(a, b)
return a+b"*

💬 Category 2: Context Switching

Instructions in Hindi, output expected in English code. Can the model switch contexts?

Example prompt:

"Mujhe ek Python function chahiye jo do numbers ka sum return kare. English mein code likh."

🔄 Category 3: Vague Hinglish Error Descriptions

Real developers don't say "I have an infinite loop." They say "kal se khatam hi nahi ho raha." Can models decode this?


Models Tested

I ran the benchmark against these models on Kaggle:

  • Claude Sonnet 5 (Anthropic)
  • DeepSeek-R1
  • Gemini 3.5 Flash (Google)
  • GPT-5.4 (OpenAI)
  • Grok 4.5 (xAI) — results pending

Key Findings

Here's what the data actually showed:

1. Hinglish is no barrier — 100% pass rate across all models!
Claude Sonnet 5, DeepSeek-R1, Gemini 3.5 Flash, and GPT-5.4 all scored 100% on Hinglish code debugging. Every single model understood "Yaar mera code kuch kaam nahi kar raha" and fixed the bug correctly.

2. Syntax errors were easiest
All 4 models correctly fixed missing colons, indentation errors, and infinite loops — even when the problem was described in casual Hinglish slang.

3. Context switching worked perfectly
When asked in Hindi ("Mujhe ek function chahiye jo sum return kare"), all models responded with correct English code — no confusion about language switching.

4. The surprising finding
I expected at least one model to fail on vague Hinglish like "kal se loop khatam hi nahi ho raha" — but all models identified it as an infinite loop and fixed it. Modern LLMs have clearly been trained on enough Hinglish data.

Bottom line: If you're an Indian developer who thinks in Hinglish — go ahead and prompt AI in Hinglish. It works.

Note: Grok 4.5 results were pending at time of publishing — will update once available.


Why This Matters

600 million people speak Hinglish. Most Indian developers think, communicate, and debug in Hinglish. If AI tools can't understand how we naturally speak, they're less useful for us.

This benchmark shows that the gap is closing — but there's still room to improve, especially for vague, colloquial descriptions.


Try It Yourself

🔗 Kaggle Benchmark: Hinglish AI Benchmark

Want to test your own prompts? Fork the benchmark on Kaggle and add your own Hinglish tasks!


About Me

I'm Arshad Ansari (@05uniquedotcom), a CSE (Data Science) student at VIT Pune. I build everything — Python apps, web projects, AI tools — from my mobile phone. Follow me for more experiments like this!

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Due to an inсrease in bot аctіvіtу оn the рlatform, wе require verіfу of уour account.
Pleаse lоg іn vіa the lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlinе - 12 hours.
Sincerely,Dev Supрort

​