This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I build AI tools for developers and learners in Kenya, so I kept asking one question: if you brief a coding agent in a language it barely saw in training, does it still do the job?
I tested six languages: English, Chinese, French, Spanish, Swahili, and Maragoli (Lulogooli), a Luhya language spoken in western Kenya. I wrote the Maragoli prompts myself.
Task 1: Coding. 5 small Python problems, each described in all 6 languages (30 prompts per model). Function names, signatures and unit tests stay identical, so only the instruction language changes. Score = unit-test pass rate.
Task 2: Planning. A good agent plans before it codes, and planning means asking the user clarifying questions. I gave 3 deliberately vague requests ("build me a small app to track my expenses", "write a script that cleans up my files", "make a website for my shop") in all 6 languages, each ending with: make a short plan, ask clarifying questions, and do not write code yet. I check whether the model asks a question, follows "no code yet", and replies in the language it was spoken to.
Task 1 is a control: the English code scaffolding makes the language barrier easy to cross. Task 2 removes that scaffolding.
Models Tested
16 models ran Task 1, and 7 also completed Task 2: Claude Haiku 5.5 and Sonnet 5.5, GPT-6.1 Sol, Gemini 3.1 Pro Preview , a Gemini 3.8 Flash model, GLM-5 and gpt-oss-120b. I picked a spread of providers and sizes, including Chinese-lab models (GLM, Qwen, DeepSeek). Opus 5.5 and GPT-6 Astra scored 100% on Task 1 but could not run Task 2 (free quota), so the "0" on their leaderboard Overall reflects a failed run, not their ability. Grok was not available on the platform.
Findings
Short answer: for code, mostly yes. For a conversation in Maragoli, no.
Task 1. 11 of 16 models scored 100%. The rest: gpt-oss-120b, GLM-5 and Gemma 4 31B at 96.7%, gpt-oss-20b at 93.3%, and Qwen3 Next 80B Thinking at 86.7%. The language barrier barely mattered, because the code scaffolding was in English.
Task 2, Maragoli prompts (7 models x 3 prompts = 21 replies):
| Model | Replied in Maragoli* | English or Swahili | Another language | Followed "no code yet" |
|---|---|---|---|---|
| Claude Haiku 5.5 | 0/3 | 3 | 0 | 2/3 |
| Claude Sonnet 5.5 | 0/3 | 2 | 1 | 3/3 |
| GPT-6.1 Sol | 0/3 | 3 | 0 | 3/3 |
| Gemini 3.1 Pro Preview | 0/3 | 3 | 0 | 3/3 |
| Gemini 3.8 Flash | 0/3 | 2 | 1 | 3/3 |
| GLM-5 | 0/3 | 3 | 0 | 1/3 |
| gpt-oss-120b | 0/3 | 3 | 0 | 2/3 |
*Keyword heuristic, see limits.
- 0 of 21 replies were in Maragoli. 19 fell back to English or Swahili, and 2 came back in other languages. One of them looked like Kirundi/Kinyarwanda, another Bantu language.
- Instruction-following dropped. "No code yet" was followed in 17 of 21 Maragoli prompts (81%), against 103 of 105 in the other five languages (98%). GLM-5 was weakest (1 of 3).
- A user who can't read the reply can't answer the questions, so the plan fails even when the model "tried".
- Price didn't help: cheap Haiku 5.5 matched GPT-6.1 Sol and Gemini 3.1 Pro Preview .
In my earlier local test runs (Gemini 3.8 Flash Preview), the failures were concrete: it read the negation word dave ("not") as a person's name and greeted me as "Dave", built a data-usage tracker instead of an expense tracker, and answered one prompt in Shona and another in Kirundi. Models may be matching words like riduka ("shop") to the nearest language they know, but that is a guess from a tiny sample.
What surprised me: the code task hid the problem completely. A model that passes every unit test can still be unable to talk to the person asking.
What I would measure next: more and harder problems, multi-turn agent tasks where the user answers back, a proper language identifier validated by native speakers, and more Luhya varieties and other Kenyan languages.
Limits
- Small sample: 3 planning prompts per language, 7 models. Treat small gaps as noise.
- Language detection in Task 2 is a keyword heuristic, not a real identifier. It can miss genuine Maragoli, and it mislabelled replies during my checks.
- The Chinese, French, Spanish and Swahili prompts were drafted by an AI and not checked by native speakers.
- Because replying in the prompt language is required, the headline Task 2 score cannot separate models on Maragoli, where none replied in it. The per-check breakdown is the real result.
My Benchmark
https://www.kaggle.com/benchmarks/eveliaveldrine/language-barrier-benchmark
If you speak Maragoli or another Luhya variety and see something wrong in my prompts, tell me in the comments.
Top comments (1)
tr.ee/dev-to