DEV Community

Cover image for TeluguMixBench: Evaluating LLMs on Telugu and Telugu-English Code-Mixed Tasks
Poojitha Saranya
Poojitha Saranya

Posted on

TeluguMixBench: Evaluating LLMs on Telugu and Telugu-English Code-Mixed Tasks

What I Benchmarked

Can large language models reliably understand Telugu and Telugu-English code-mixed prompts?

To explore this, I created TeluguMixBench, a small benchmark containing 20 test cases across five categories:

  • Reading comprehension: Understanding Telugu passages and answering questions.
  • Reasoning: Solving simple problems expressed in Telugu.
  • Translation: Translating Telugu sentences into English.
  • Code-mixed understanding: Handling prompts that combine Telugu and English.
  • Instruction following: Following constraints such as providing a specified number of points or answering in a requested language.

I chose this problem because multilingual AI evaluation should consider not only widely used languages but also regional languages and the way people naturally mix languages in everyday conversations.

Models Tested

I used Kaggle Benchmarks to evaluate models on TeluguMixBench.

The leaderboard included:

  • Claude Opus 4.5
  • Claude Haiku 4.5
  • Claude Haiku 5.5
  • Gemini 3.7 Flash

I wanted to explore how models from different providers performed on the same set of language tasks and how their scores compared.

Findings

The initial leaderboard showed the following scores:

Model Score
Claude Opus 4.5 100%
Claude Haiku 4.5 100%
Claude Haiku 5.5 100%
Gemini 3.7 Flash 95%

These results are an interesting starting point, but they should be interpreted cautiously.

The benchmark currently contains only 20 examples, and some evaluation checks rely on expected-answer matching rather than a full assessment of meaning. A model may include an expected keyword without providing a completely correct answer. Perfect scores therefore do not establish that a model reliably understands Telugu across different contexts.

My main takeaway is that the quality of an evaluation method matters as much as the scores it produces. A useful next step is to expand the dataset, improve semantic evaluation, test more varied Telugu-English prompts, and examine incorrect responses individually.

My Benchmark

Explore TeluguMixBench on Kaggle:

https://www.kaggle.com/benchmarks/poojithasaranya/telugumixbench/leaderboard

The benchmark is an initial step toward evaluating regional-language and code-mixed capabilities more systematically. I hope to improve its coverage and scoring methodology in future iterations.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to