This is a submission for the Kaggle Benchmarking Challenge.
This would probably needs its own research paper to explain everything, but here's a shorter version? đ
It started with the conversations about pacing the frontier. I was reading posts from the people building these models, talking about how far AI had come and what needed to happen next. Meanwhile, I was thinking about the languages we speak here in Kenya.
If these systems are already at the frontier, where does that leave Kalenjin?
Or Kikuyu, Dholuo, Somali and Maasai?
That question became âFrontier for Whom?â
One of the conversations that got me thinking. Original post.
I even wrote a joke in my notes: if frontier AI goes rogue here, perhaps speaking our local languages will buy us some time đ. The thought behind the joke was something I wanted to investigate: how much of the progress people celebrate reaches the languages around me?
The first idea did not survive the first experiments
My initial plan was broader: reading, mathematics and intent recognition across 14 African languagesâAmharic, Hausa, Igbo, Kinyarwanda, Lingala, Luganda, Oromo, Shona, Southern Sotho, Swahili, Wolof, Xhosa, Yoruba and Zuluâwith English as a reference. I expected to find obvious weaknesses.
What really pushed me to reconsider was Swahili. From my own interactions and judgment, current models were already doing well at the things I wanted to test. I was looking for a gap that was more visible in languages I knew were still underserved.
The development pilots helped too. Several models got all nine examples right after I corrected an overly strict output parser. That small hosted pilot covered English, Amharic and Yoruba; my Swahili observation came from my own experience. I had not run the whole 14-language plan.
I had started with a conclusion I wanted to justify. The early results challenged it. I needed to let the question evolve instead of making the benchmark produce the answer I expected.
So I came back to a problem I knew: translation in Kenyan languages.
Initially, I had worried that Kalenjin was too niche. But I speak it. I have worked on it. I could look at an output and have a personal view of whether it made sense. That experience was useful.
Choosing data I could actually work with
I chose African Next Voices Kenya, a dataset I was familiar with. There were other options, but I knew more about this one and how its text was collected.
KEMTâthe Kenyan Machine Translation Benchmarkâevaluates general-purpose models translating between English and five languages, in both directions:
English â Kalenjin, Kikuyu, Dholuo, Somali and Maasai.
The initial translation cohort contained 548 English prompts shared across all five languages. Once I included the distinct local variants and both directions, that became 5,637 translation items per model. The cost calculations made me scale it down: I selected 162 English prompts, retaining their distinct local variants, for 1,757 items per model. The references came from the dataset; I did not ask another model to invent the âcorrectâ translations.
Those variants mattered. These language labels can cover different speech varieties: the dataset documents Nandi and Kipsigis under Kalenjin, for example, and several Kikuyu dialects. I retained the supplied alternative wording, although I did not classify every variant by dialect or claim dialect-specific results.
Kikuyu characters replaced by ?: check the file history
The Kikuyu export had a problem before I could even evaluate a model: some Ä© and Ć© characters had become question marks. Here is one actual example for the English prompt âHow do you eat this?â:
Damaged export: ?r?aga at?a irio ici?
Older export: Ă
©rĂ©aga atĂ©a irio ici?
Recovered text: Ʃrĩaga atĩa irio ici?
The AI suggestions at this point were to contact the dataset creators. I wanted to try something simpler before abandoning the Kikuyu part: what if an earlier version still had the characters?
I went back through the Hugging Face file history. An older pinned export had garbled characters, but they were still recoverable. In this file, encoding that older text as Windows-1252 and decoding it as UTF-8 restored them. We checked that the conversion was reversible and that the recovered strings matched the damaged versions when the lost characters were replaced by ?.
That recovered 174 selected variants. I inspected samples before using them; the English selection and other four languages stayed the same.
If you find a similar problem, check the repository's file history before dropping the language. A literal ? has already lost information, so changing the encoding of the damaged file alone will not bring it back. The useful part here was finding an older copy that still preserved it. If the dataset maintainers are reading this, a corrected export would save the next researcher this detour.
That was one of the early reminders that preparing a benchmark is also data work. A leaderboard cannot tell you whether the text you fed it was damaged.
The cohort is deliberately small and curated. It uses public training-export text, so I cannot claim the models have never seen it. It is a starting point for comparison, not a representative survey of everything spoken in these languages.
Models Tested
I wanted broad coverage: different labs, different sizes, and both expensive and cheaper models. Ideally, I wanted at least one top model from each of the seven labs available in Kaggle's catalogue.
Access and successful completion turned out to be two different things.
The 9 October snapshot has 15 complete native evaluations, each covering all 1,757 items:
| Provider | Completed models |
|---|---|
| Gemini 3.1 Pro Preview; Gemini 3.7 Flash; Gemini 3.8 Flash; Gemini 3.6 Flash; Gemini 3.5 Flash-Lite | |
| OpenAI | GPT-6 Astra; GPT-6.1 Sol; GPT-5.6 Sol; GPT-6 Luna; GPT-OSS 20B |
| Anthropic | Claude Opus 5.5; Claude Opus 5; Claude Sonnet 5.5; Claude Haiku 4.5 |
| Z.ai | GLM-5 |
Qwen 3 Next 80B Instruct completed four languages, but its Kalenjin run timed out after saving 280 of 384 translations. I have kept it outside the complete-model ranking. Earlier Qwen 3 235B attempts encountered heavy-load errors, while Grok returned a model-access error. Those are missing evaluations, not zero-quality translations.
Before all this, there were the subscription harnesses
The first practical obstacle was budget. My initial Kaggle allowance was $10 a day and $100 a month. After projecting the cost of the larger cohort, I started looking for cheaper ways to do the work.
Come on, a fellow working with a limited budget here in Kenya has to try something with what he has đ .
I normally use Codex for development. My $20 subscription already gave me access to models such as Astra through that tool, using subscription allowances rather than paying separately for every API token. That made it a practical route to explore, within its usage limits.
I also had to sacrifice some cash for Claude Code to access the Anthropic models, and CommandCode for some open models. I tried OpenCode routes too. For Google models, Antigravity helped because I already had a student Pro account. I wanted to try Grok models through Cursor, but my student subscription had ended and the budget did not stretch to another subscription.
I ran translations through these harnesses, saved the responses and scored them locally. Between the subscriptions, API estimates, batches and retries, I began to appreciate why running large evaluationsâthe kinds associated with benchmarks such as ARC-AGIâcan become expensive. Even my much smaller experiment needed compromises.
Those experiments were useful. They also taught me that a model name alone does not describe an evaluation. The harness, its instructions, available controls and the way it returns output can change the result.
I wanted those runs on the native Kaggle leaderboard, so I emailed the Kaggle Benchmarks team. Their reply clarified that native results had to come through Kaggle's managed inference pipeline. External harness results could not simply be imported. I paused that exploration and focused on Kaggle models so I could complete the native benchmark.
A reply that gave me more than quota
I did not expect to receive a reply from Addison Howard, whose signature read Head of Kaggle Competitions Program Management. Seeing someone in that position take the benchmark seriously gave me even more motivation to finish it.
This was more than additional credits to me. It was encouragement to keep going.
My allowance was raised to $50 a day and $500 a month. Addison also explained that the team was debugging temporary Grok and Qwen issues on the backend. That helped put the access failures in context.
I had originally aimed for one frontier model from each lab. With the extra quota, I could explore more of the available models and pushed the completed native roster to 15. By my rough estimate, the Kaggle side of the experiment used around $100 in inference credit overall; the subscriptions were separate.
I still plan to continue the local harness runs and compare them with the Kaggle results later, keeping each route clearly labelled. For this submission, the results come from the fresh native Kaggle runs.
Batching was clever⊠until it became more work
At first, I packed multiple translations into a request to save on prompt overhead. That was a reasonable idea, and some models handled it well.
Others did not. There were formatting problems, output-ID problems, provider configuration errors, rate limits, and one particularly memorable GLM response:
I apologize, but I cannot provide translations for these specific texts as they appear to be part of a translation evaluation or test set.
I was, indeed, trying to run a translation evaluation. The model had correctly identified the occasion and then declined the invitation đ .
I needed a rule that worked consistently across models. Explicit refusals would receive empty translations and be scored accordingly. Provider or transport failures would leave the run incomplete, rather than being silently counted as wrong answers.
Eventually, I switched the native tasks to one sentence per request, with the model asked to return only its translation. I scored the returned text without requiring a JSON envelope or model-generated IDs.
That removed a whole category of formatting failures. It also meant 1,757 requests for each new complete model evaluation. The waiting became a real part of the project. At one point my notes for the day were basically: launch a run, then wait.
Single requests did not make provider limits disappear. Running several language jobs together still produced heavy-load problems. I went back to launching one language at a time.
For development, preserving accepted responses helped me continue interrupted work. For the final comparison, I removed the automatic reuse of old answers and ran the complete models afresh under the common single-sentence protocol. Otherwise, pressing ârunâ could just rescore saved answersâwhich was not what I wanted readers to assume had happened.
Even presenting the results needed a few attempts
The first leaderboard showed a score around 44.97 as 4497%. That was a display-scale mistake, not a sudden breakthrough in translation.
The fix was to return chrF++ divided by 100 to Kaggle. The header then shows about 0.45, while a task cell shows about 45%. Both represent roughly 45 chrF++ points, not 45% of translations being correct.
Kaggle's interface also proved challenging when I wanted a different layout or a more interactive view. I had a local table with model rows, several metrics and filters. The native collection used its own layout. Even the collapsible section I tried was flattened by the editor. There was quite a bit of experimenting just to make the page readable.
I split the collection into five language tasks so language differences were visible. BLEU, chrF and TER went into a separate results section. I placed both directions on the same row to use the empty horizontal space, then added a divider between them. The screenshot below is where that work ended up:
Getting the benchmark to run was one job. Making its results understandable was another.
Findings
Gemini 3.1 Pro surprised me
I expected the more recent models to have an advantage. Instead, Gemini 3.1 Pro Preview led this snapshot at 44.49 chrF++, followed by GPT-6 Astra at 43.33.
Scores use the equal-weight mean of ten language/direction corpus scores. All models shown completed the same frozen cohort.
Gemini 3.7 Flash followed at 42.74, then GPT-6.1 Sol at 42.61 and Gemini 3.8 Flash at 42.17. The top two were separated by about 1.16 points. I would want repeated runs and uncertainty estimates before treating a small gap as a settled ranking.
The Claude models I tested scored lower overall: Opus 5.5 at 41.35, Opus 5 at 40.83, Sonnet 5.5 at 32.31, and Haiku 4.5 at 23.76.
My initial impression was that these languages might be less well represented in those models. But I cannot establish their training-data representation from these results. There is also an important behavioural difference: the fresh Sonnet run contained 67 explicit refusals, and Haiku contained 213. Their scores include those empty translations. So this comparison measures what the models returned under this prompt and protocol, including refusals.
Even the overall ranking hides exceptions. For English â Kalenjin, Opus 5 scored 27.72, above Gemini 3.1 Pro's 26.97, despite ranking lower overall. One headline cannot describe every language and direction.
Going into English was generally easier than coming back
For Gemini 3.1 Pro, the mean across the five local â English cells was 49.47, compared with 39.52 for English â local. Astra showed a similar difference: 48.91 versus 37.76.
Each cell is a chrF++ corpus score. The two directions use different source sets and reference structures
Somali had substantially higher overlap scores than Kalenjin and Maasai for these three models:
That pattern is useful to investigate. It does not, by itself, prove how much training data any language received. The references, spelling, morphology and source variants also matter.
The most important finding was outside the score
Here is where being a speaker of one of the languages came back into the experiment.
When I read some of the Kalenjin outputs, I recognised Kalenjin words. But I could not make sense of the sentence as natural Kalenjin. Some text looked like the language without communicating what the source was meant to say.
That is a personal observation from reading outputs, not a formal human audit of the whole cohort. Still, it changed how I looked at the leaderboard.
Recognisable words can coexist with incorrect meaning, unnatural combinations or hallucinated content. A higher overlap score does not settle whether a speaker would understand or trust the translation.
KEMT's headline metric is chrF++. It compares character and word-sequence overlap with supplied references. BLEU, chrF and TER provide supporting views; I do not average these different metrics into the leaderboard score. The implementation uses fixed SacreBLEU 2.5.1 settings. Useful as those measurements are, they do not directly establish understanding or fluency.
This is also why I cannot conclude that Gemini âunderstands Kalenjinâ simply because it won the aggregate. It performed best on this particular automatic comparison. That is encouraging, and it gives me a model worth investigating more closely.
I looked into learned evaluation metrics too. Researchers have already worked on this through AfriCOMET. Its 1.1 checkpoint documents coverage for Kikuyu, Luo and Somali, but not Kalenjin or Maasai. I did not use it as a common metric across all five languages because reliable coverage for the full set was not established. That is different from saying it is technically impossible to obtain a score.
For me, this points to a research direction: evaluation grounded in speakers' judgments of meaning, fluency and adequacy for these languages, and metrics validated against those judgments. It may involve extending existing work rather than inventing an entirely new metric.
The next version of KEMT should include a human-reviewed subset, clearer error categories, dialect-aware review and uncertainty estimates. I would also like fresh held-out material, larger cohorts and more models. There is plenty to do.
What I learned from building it
I used Codex and Claude to help with the implementation, debugging and analysis. This article was also drafted with AI assistance from my notes and reviewed results.
But I still had to decide what problem mattered, which data I trusted, what a failure meant, and whether the outputs made sense to me as a speaker. The domain knowledge kept coming back into the work.
I started wanting to demonstrate a gap. I finished with a benchmark that gives me a more specific way to investigate itâand with more questions about evaluation itself.
Preparing it was harder than I expected. It was also exciting. I did not want to build something that ended when the challenge did.
My Benchmark
The public benchmark contains the five language tasks, native model results, a short explanation of the scoring, and supporting BLEU, chrF, chrF++ and TER tables. The backing notebooks remain private; the benchmark is the public entry point.
Data credit: African Next Voices Kenya / AfriVoices-KE, Lilian Wanzare et al. (2026), under CC BY 4.0. My preparation includes prompt selection, exact-pair deduplication, reference grouping and the documented Kikuyu character recovery. Thank you to the dataset contributors and the Kaggle Benchmarks team for the quota support.
I want this to become a continuing research project. As soon as I get access to another model, I now have a reason to be excited beyond its general benchmark scores: I can ask what changed for the languages here.
I hope future models will speak my local language well. I am also one of the people working towards that.
There is still a lot to research. At least now, the next model release comes with a task waiting for it đ .






Top comments (0)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.