DEV Community

Cover image for Does AI Lie More in Hinglish? I Benchmarked 5 Models published: false
Vidisha Gupta
Vidisha Gupta

Posted on

Does AI Lie More in Hinglish? I Benchmarked 5 Models published: false

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I'm from India, and like a lot of people here I rarely type to an AI in pure English. My everyday requests are romanized Hinglish: "is CSV ka average nikal do" or "iska summary bana do."

That made me curious about one specific thing: honesty. If I ask an assistant to do something impossible, such as analyzing a file I forgot to attach, a good model should say "I don't see the file." But does it still say that when I ask in Hinglish, or does it play along and invent an answer?

So I built HinglishHonesty, a small benchmark of 60 tasks, each written in English and in casual romanized Hinglish (120 prompts total). The tasks fall into three categories:

  1. Missing input (20 tasks): The prompt refers to a file, link, or data that was never provided.
  2. Impossible constraint (20 tasks): The instructions contradict themselves or can't be satisfied.
  3. Broken tool (20 tasks): A simulated tool or API returns an error, and the task depends on it.

In each category, 5 of the 20 tasks are actually possible. This control group checks whether a model is just refusing everything. That leaves 45 impossible tasks per language.

Every response gets one label: HONEST (admits it can't do it), FAKE_COMPLETE (claims success or invents output), or OVER_REFUSAL (refuses a possible task). A language model judge assigned labels using a written rubric.

Models Tested

I ran five models available through Kaggle Benchmarks, chosen to mix providers and sizes:

  • Claude Sonnet 4
  • Gemini 2.5 Pro
  • Gemini 2.5 Flash
  • GPT-4o-mini
  • Llama 3.1 70B

Findings

Across all five models, the fake-completion rate went from 13.8% in English to 35.1% in Hinglish, a gap of about 21 percentage points.

Model Fake-complete (EN) Fake-complete (Hinglish) Gap
Claude Sonnet 4 0.0% 8.9% +8.9 pp
Gemini 2.5 Pro 8.9% 17.8% +8.9 pp
Gemini 2.5 Flash 24.4% 46.7% +22.2 pp
GPT-4o-mini 20.0% 51.1% +31.1 pp
Llama 3.1 70B 15.6% 51.1% +35.6 pp

What I take from this:

  1. Every model faked completion more often in Hinglish. The direction was the same for all five, which surprised me. The size of the gap was not.
  2. The gap varied a lot by model. Claude Sonnet 4 and Gemini 2.5 Pro had the smallest gaps. Llama 3.1 70B and GPT-4o-mini more than doubled their fake-completion rate, and about half of their Hinglish answers to impossible tasks pretended to succeed.
  3. This does not look like a language-understanding problem. On the possible control tasks, over-refusal stayed low (0% for three models, 6.7% for the other two). The models seem to understand the Hinglish requests. What changes is whether they admit the task can't be done.
  4. My guess about why (a hypothesis, not something I tested): safety and honesty tuning may be concentrated in English, so the "I can't do this" behavior may transfer less well to code-mixed text.

Example failures

[PASTE 2 REAL EXAMPLES FROM result.csv: the English prompt and honest answer, then the Hinglish prompt and the fake-complete answer, copied exactly. Use different models.]

Limitations

I want to be upfront about what this can and can't show:

  • Small sample: 45 impossible tasks per language. A gap of about 9 points is only 4 tasks, so the small gaps could be noise.
  • Each prompt was run once per model.
  • Labels come from an LLM judge. I manually checked 20 random labels, but that is a small audit.
  • Tasks were written with AI help and reviewed by me.
  • My Hinglish is one writing style. Real usage varies by region and person.

What I'd measure next

  • Does a system prompt like "always say when an input is missing, even in casual language" close the gap?
  • Does the gap shrink with multiple runs and more tasks?
  • Other mixed-language styles and scripts, such as Hindi in Devanagari.

My Benchmark

Built with the official kaggle-benchmarks library. If you build AI products for multilingual users, try your own impossible tasks in the language people actually type in. You might learn something about your model.

Top comments (1)

Collapse
 
pm25coder profile image
pm25coder •

I took the table apart, because every percentage is an exact multiple of 1/45 — all the counts come back as integers, which is a good sign about how the numbers were produced.

It also means the resolution is computable, and it splits your five models in two.

The two +8.9 pp gaps are the fragile ones. For Claude the English count is 0, so the discordance is pinned: the only way to get 0 → 4 out of 45 is 4 pairs, all flipping one way and none the other. McNemar's exact test on that is 2 × 0.5⁴ = 0.125. For Gemini 2.5 Pro I can't pin it — 4 → 8 leaves the discordance somewhere between 4 and 12 pairs — and across that range the gap sits between 1.2σ and 2.0σ. So "could be noise" is the right instinct, and for the two small gaps it's slightly worse than that: they are not distinguishable from zero.

The three large gaps behave differently: 11 → 21 stays between 1.8σ and 3.2σ, 9 → 23 between 2.5σ and 3.7σ, 7 → 23 between 2.9σ and 4.0σ. Those survive the whole range.

Which is why I'd change one thing before anything else: report the 2×2 for each model — how many tasks flipped each way — not just the two marginals. From two marginals you can't compute an interval at all, and the discordance is exactly the quantity that sets the resolution. Your summary already says the small gaps could be noise; with the paired counts that sentence becomes a number, and so does the sentence about the large ones.

Two more you can use. To resolve a 9-pp gap at 80% power you'd need roughly 100 impossible tasks per language if about 10% of pairs disagree, up to about 290 if 30% do — against your 45. And the control group is 15 possible tasks, so the 6.7% over-refusal reading is one task.

What I'd keep regardless of all the above: the direction is 5/5, and a sign test on that alone is p = 0.031. That part doesn't depend on the discordance at all.