What I Benchmarked
Can large language models reliably understand Telugu and Telugu-English code-mixed prompts?
To explore this, I created TeluguMixBench, a small benchmark containing 20 test cases across five categories:
- Reading comprehension: Understanding Telugu passages and answering questions.
- Reasoning: Solving simple problems expressed in Telugu.
- Translation: Translating Telugu sentences into English.
- Code-mixed understanding: Handling prompts that combine Telugu and English.
- Instruction following: Following constraints such as providing a specified number of points or answering in a requested language.
I chose this problem because multilingual AI evaluation should consider not only widely used languages but also regional languages and the way people naturally mix languages in everyday conversations.
Models Tested
I used Kaggle Benchmarks to evaluate models on TeluguMixBench.
The leaderboard included:
- Claude Opus 4.5
- Claude Haiku 4.5
- Claude Haiku 5.5
- Gemini 3.7 Flash
I wanted to explore how models from different providers performed on the same set of language tasks and how their scores compared.
Findings
The initial leaderboard showed the following scores:
| Model | Score |
|---|---|
| Claude Opus 4.5 | 100% |
| Claude Haiku 4.5 | 100% |
| Claude Haiku 5.5 | 100% |
| Gemini 3.7 Flash | 95% |
These results are an interesting starting point, but they should be interpreted cautiously.
The benchmark currently contains only 20 examples, and some evaluation checks rely on expected-answer matching rather than a full assessment of meaning. A model may include an expected keyword without providing a completely correct answer. Perfect scores therefore do not establish that a model reliably understands Telugu across different contexts.
My main takeaway is that the quality of an evaluation method matters as much as the scores it produces. A useful next step is to expand the dataset, improve semantic evaluation, test more varied Telugu-English prompts, and examine incorrect responses individually.
My Benchmark
Explore TeluguMixBench on Kaggle:
https://www.kaggle.com/benchmarks/poojithasaranya/telugumixbench/leaderboard
The benchmark is an initial step toward evaluating regional-language and code-mixed capabilities more systematically. I hope to improve its coverage and scoring methodology in future iterations.
Top comments (1)
tr.ee/dev-to