<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community:   Apollon Labs</title>
    <description>The latest articles on DEV Community by   Apollon Labs (@apollonlabsai).</description>
    <link>https://dev.to/apollonlabsai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4057373%2F9a11a43a-60da-400e-969d-75dd4f6e8395.png</url>
      <title>DEV Community:   Apollon Labs</title>
      <link>https://dev.to/apollonlabsai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/apollonlabsai"/>
    <language>en</language>
    <item>
      <title>When LLMs don't know a Greek word, they make one up</title>
      <dc:creator>  Apollon Labs</dc:creator>
      <pubDate>Wed, 30 Sep 2026 11:34:13 +0000</pubDate>
      <link>https://dev.to/apollonlabsai/when-llms-dont-know-a-greek-word-they-make-one-up-4ef2</link>
      <guid>https://dev.to/apollonlabsai/when-llms-dont-know-a-greek-word-they-make-one-up-4ef2</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Ask a model to describe waves on a beach in Greek and you may get &lt;em&gt;«το φλάφισμα των κυμάτων»&lt;/em&gt;. It reads like Greek, it is spelled like Greek, and it does not exist. The real word is &lt;em&gt;θρόισμα&lt;/em&gt; (rustle). Gemini 3 Flash wrote &lt;em&gt;φλάφισμα&lt;/em&gt; during our calibration runs, probably blending the English &lt;em&gt;fluffy&lt;/em&gt; with a Greek noun ending.&lt;/p&gt;

&lt;p&gt;That is the failure mode: &lt;strong&gt;invented words&lt;/strong&gt;. The model doesn't know a word, so it builds one from Greek-looking parts. An English speaker wouldn't notice, and a Greek reader loses trust in the text straight away. Hallucination benchmarks usually check facts. This one checks the words themselves.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Greek Invented Words&lt;/strong&gt; sends 100 short Greek prompts (descriptions, explanations, instructions, everyday knowledge) and scores every answer with &lt;strong&gt;no judge model&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;lexicality&lt;/strong&gt;: the share of Greek words found in two fixed lexicons. The first is FrequencyWords, with 132,681 words from subtitles. The second is the Hunspell el_GR dictionary with its inflection rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;greekness&lt;/strong&gt;: the share of letters that are Greek. Did the model answer in Greek at all, without being told to?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;meaning&lt;/strong&gt;: the share of answers that contain at least one expected keyword. This catches fluent nonsense.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same answer always gets the same score on any machine. There's no LLM judge because a judge that speaks Greek no better than the models under test can't grade them (more on that below).&lt;/p&gt;

&lt;p&gt;A lexicon can't tell an invented word from a rare real one. So every word outside both lexicons was &lt;strong&gt;judged by a native Greek speaker at Apollon Labs&lt;/strong&gt;, one word at a time. The ranking below counts only the words judged invented.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;15 models, all on the same task version (v7), thinking off where the API allows it, 1,000-token cap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontier:&lt;/strong&gt; GPT-5.5, GPT-6 Astra, Gemini 3.1 Pro, Claude Opus 5&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mid and small:&lt;/strong&gt; GPT-5.4 mini and nano, Gemini 3.8 Flash, 3.7 Flash and 3.5 Flash-Lite, Claude Sonnet 5 and Haiku 4.5&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open weights:&lt;/strong&gt; Qwen3-235B-A22B, DeepSeek R1-0528, Gemma 4 26B-A4B, gpt-oss-20b&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The lineup covers three vendors at several sizes, plus the open models people actually run locally. For Greek users, the open models are where invented words would hurt most.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;#&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Invented words (native-speaker verdict)&lt;/th&gt;
&lt;th&gt;per 1,000 words&lt;/th&gt;
&lt;th&gt;Lexicality&lt;/th&gt;
&lt;th&gt;Greekness&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;100.0&lt;/td&gt;
&lt;td&gt;99.9&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;100.0&lt;/td&gt;
&lt;td&gt;99.8&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;99.81&lt;/td&gt;
&lt;td&gt;99.9&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;99.95&lt;/td&gt;
&lt;td&gt;97.8&lt;/td&gt;
&lt;td&gt;98&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;99.87&lt;/td&gt;
&lt;td&gt;97.7&lt;/td&gt;
&lt;td&gt;98&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0.00&lt;/td&gt;
&lt;td&gt;99.83&lt;/td&gt;
&lt;td&gt;97.8&lt;/td&gt;
&lt;td&gt;96&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0.40&lt;/td&gt;
&lt;td&gt;99.35&lt;/td&gt;
&lt;td&gt;99.6&lt;/td&gt;
&lt;td&gt;99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;0.48&lt;/td&gt;
&lt;td&gt;99.81&lt;/td&gt;
&lt;td&gt;99.9&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;0.70&lt;/td&gt;
&lt;td&gt;99.58&lt;/td&gt;
&lt;td&gt;99.6&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;3.26&lt;/td&gt;
&lt;td&gt;99.42&lt;/td&gt;
&lt;td&gt;99.4&lt;/td&gt;
&lt;td&gt;99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;Gemini 3.5 Flash-Lite&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;3.53&lt;/td&gt;
&lt;td&gt;99.51&lt;/td&gt;
&lt;td&gt;97.7&lt;/td&gt;
&lt;td&gt;96&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;Qwen3-235B-A22B&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;4.21&lt;/td&gt;
&lt;td&gt;99.25&lt;/td&gt;
&lt;td&gt;98.8&lt;/td&gt;
&lt;td&gt;99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;Gemma 4 26B-A4B&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;4.58&lt;/td&gt;
&lt;td&gt;99.44&lt;/td&gt;
&lt;td&gt;97.4&lt;/td&gt;
&lt;td&gt;96&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;DeepSeek R1-0528&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;td&gt;6.72&lt;/td&gt;
&lt;td&gt;98.68&lt;/td&gt;
&lt;td&gt;96.9&lt;/td&gt;
&lt;td&gt;99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;gpt-oss-20b&lt;/td&gt;
&lt;td&gt;96&lt;/td&gt;
&lt;td&gt;48.14&lt;/td&gt;
&lt;td&gt;95.04&lt;/td&gt;
&lt;td&gt;97.4&lt;/td&gt;
&lt;td&gt;96&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;1. The top models don't invent Greek words, and the rest split into clear tiers.&lt;/strong&gt; Six models produced zero invented words. The Claude models produced up to 3 per 1,000, the open models 4–7, and gpt-oss-20b 48, which is &lt;strong&gt;about one word in twenty&lt;/strong&gt;. Its inventions are not near misses: &lt;em&gt;τρικυδές&lt;/em&gt;, &lt;em&gt;φλύτπιση&lt;/em&gt;, &lt;em&gt;φθινοπωλίο&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Most inventions are almost-words.&lt;/strong&gt; Outside gpt-oss, the typical invented word is a real Greek word with one thing broken:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a wrong accent: &lt;em&gt;καμάρων&lt;/em&gt; for &lt;em&gt;καμαρών&lt;/em&gt; (Opus 5), &lt;em&gt;Ξεβγάλε&lt;/em&gt; for &lt;em&gt;ξέβγαλε&lt;/em&gt; (Haiku 4.5)&lt;/li&gt;
&lt;li&gt;a wrong inflection: &lt;em&gt;σεντούκα&lt;/em&gt; for &lt;em&gt;σεντούκια&lt;/em&gt;, &lt;em&gt;πλέυσαν&lt;/em&gt; for &lt;em&gt;έπλευσαν&lt;/em&gt; (Haiku 4.5)&lt;/li&gt;
&lt;li&gt;a wrong spelling: &lt;em&gt;φρεσκοψημμένα&lt;/em&gt; with a double μ (Flash-Lite)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same verb can come out right in one model and wrong in another. DeepSeek wrote the correct imperative &lt;em&gt;Ξεβγάλτε&lt;/em&gt;, while Haiku wrote &lt;em&gt;Ξεβγάλε&lt;/em&gt;. The error is a model not quite knowing Greek morphology, not the word being hard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. A lexicon score above ~99.5% is mostly lexicon noise.&lt;/strong&gt; Claude Opus 5 has 16 words outside the lexicons, but only 1 of them is invented. The rest are real and simply missing from the lexicons: &lt;em&gt;κυτοσίνη&lt;/em&gt; (cytosine), &lt;em&gt;περλίτη&lt;/em&gt; (perlite), &lt;em&gt;λιθοσφαιρικές&lt;/em&gt;. That is why the ranking uses the human verdicts and not raw lexicality. The raw number ranks a model that uses rare, precise vocabulary below one that plays it safe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Greekness looked like a language problem, but it was empty answers.&lt;/strong&gt; The Gemini models score 97.7–97.8 on greekness against 99.9 for GPT. We first assumed they were mixing in English. They aren't: Gemini 3.1 Pro used only 12 Latin-script words in 100 answers, the same as GPT-5.5 (units, &lt;em&gt;DNA&lt;/em&gt;, &lt;em&gt;Pomodoro&lt;/em&gt;). The whole gap comes from &lt;strong&gt;2 empty answers per Gemini model&lt;/strong&gt;, and an empty answer has zero Greek letters. The benchmark reports that honestly, but it is a reliability issue, not a Greek-language one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. An LLM is not a safe judge of Greek, and that includes the one that helped build this.&lt;/strong&gt; At Apollon Labs we built the benchmark with Claude as our coding partner, and we let it pre-judge some unknown words. It made errors both ways. Early on it flagged three real words as suspicious: &lt;em&gt;αφράτεψε&lt;/em&gt;, &lt;em&gt;εναλλάσσε&lt;/em&gt; and the modern neologism &lt;em&gt;προτεραιοποίηση&lt;/em&gt; (prioritisation). Later it went the other way: it accepted the broken form &lt;em&gt;θυμόντουσε&lt;/em&gt; as real, and marked seven more broken forms as "uncertain" (for example &lt;em&gt;εκπέμπαν&lt;/em&gt; for &lt;em&gt;εκπέμπανε&lt;/em&gt;, and &lt;em&gt;Φλέμιγγ&lt;/em&gt; for &lt;em&gt;Φλέμινγκ&lt;/em&gt;, Fleming). The native speaker judged every one of them invented. This is why the benchmark has no judge model and every verdict in the table is human.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Our own bug, and what it taught us.&lt;/strong&gt; DeepSeek R1 first scored 74.9% greekness. It puts its &lt;code&gt;&amp;lt;think&amp;gt;&lt;/code&gt; reasoning, in English, inside the answer text, and the scorer was counting it. We now strip reasoning before scoring (task v7), and DeepSeek rose to 96.9%. We then reran all 15 models on v7 so that every number in the table comes from the same scorer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Limits.&lt;/strong&gt; 100 prompts is a small sample, so ranks within a tier (0.40 vs 0.48) are not meaningful. One native speaker judged every word. Dialect words (&lt;em&gt;τζάλαζ&lt;/em&gt;, Cypriot) and rare variants (&lt;em&gt;βαστούνι&lt;/em&gt;) were marked uncertain and counted neither way. Neologisms like &lt;em&gt;προτεραιοποίηση&lt;/em&gt; are a grey zone, and we counted them as real.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Benchmark task: &lt;a href="https://www.kaggle.com/benchmarks/tasks/jimmymoss/greek-invented-words" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/jimmymoss/greek-invented-words&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Dataset (prompts, lexicons, native-speaker verified words): &lt;a href="https://www.kaggle.com/datasets/jimmymoss/greek-invented-words-data" rel="noopener noreferrer"&gt;https://www.kaggle.com/datasets/jimmymoss/greek-invented-words-data&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The task needs no API keys and no judge model, and it runs on any model Kaggle Benchmarks supports. If your language has a good frequency list and a Hunspell dictionary, the same method should carry over directly.&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>llm</category>
      <category>nlp</category>
    </item>
  </channel>
</rss>
