<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Soufiane Zaari</title>
    <description>The latest articles on DEV Community by Soufiane Zaari (@soufianezaari).</description>
    <link>https://dev.to/soufianezaari</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4153231%2Ff88ca42d-0f62-4f58-bb99-abbb7f138ff5.jpg</url>
      <title>DEV Community: Soufiane Zaari</title>
      <link>https://dev.to/soufianezaari</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/soufianezaari"/>
    <language>en</language>
    <item>
      <title>Darija Buddy: A Local Gemma-Powered Tutor for a Friend</title>
      <dc:creator>Soufiane Zaari</dc:creator>
      <pubDate>Sat, 03 Oct 2026 01:02:48 +0000</pubDate>
      <link>https://dev.to/soufianezaari/darija-buddy-a-local-gemma-powered-tutor-for-a-friend-1kdh</link>
      <guid>https://dev.to/soufianezaari/darija-buddy-a-local-gemma-powered-tutor-for-a-friend-1kdh</guid>
      <description>&lt;h2&gt;
  
  
  The Friend &amp;amp; The Problem
&lt;/h2&gt;

&lt;p&gt;Moving to Morocco as an international student comes with a unique linguistic hurdle: Moroccan Darija. While formal Arabic (MSA) or French might get you through official paperwork, everyday social life happens in Darija.&lt;/p&gt;

&lt;p&gt;My university classmate, Alex, has been struggling to fit into group banter and daily conversations. He understands some basics, but when people talk fast or switch between Arabizi (Latin-script Darija using numbers like 3, 7, 9) and spoken phrases, he freezes.&lt;/p&gt;

&lt;p&gt;Traditional language apps don't support Moroccan Darija well—they usually default to Modern Standard Arabic, which locals rarely speak on the street. Human tutoring is expensive, and practicing with friends can feel intimidating when you're afraid of making mistakes.&lt;/p&gt;

&lt;p&gt;He needed a safe, patient, and conversational practice partner that speaks genuine Darija, understands phonetically typed Arabizi, and explains nuances in French.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I Built: Darija Buddy 🇲🇦
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Darija Buddy&lt;/strong&gt; is a lightweight, local conversational tutor designed to help non-Moroccan beginners practice real-life Darija speech.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Capabilities:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Speaks authentic Darija:&lt;/strong&gt; It avoids rigid Modern Standard Arabic and replies in natural Moroccan expressions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multilingual input handling:&lt;/strong&gt; It comprehends Latin-script Darija (Arabizi), French, and standard Arabic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bilingual feedback loop:&lt;/strong&gt; Whenever it introduces colloquial vocabulary or idioms, it appends concise pedagogical explanations in French.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gentle error correction:&lt;/strong&gt; If the learner makes a grammar or lexical slip, Darija Buddy reformulates the phrase constructively before keeping the conversation flowing.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Technical Architecture &amp;amp; How It Works
&lt;/h2&gt;

&lt;p&gt;The whole project runs entirely offline on a personal laptop — no expensive cloud infrastructure needed.&lt;br&gt;
+-----------------------------------------------------------+&lt;br&gt;
| Local Machine |&lt;br&gt;
| |&lt;br&gt;
| +-------------------+ +--------------------+ |&lt;br&gt;
| | Gradio Web UI | &amp;lt;------&amp;gt; | Ollama Runtime | |&lt;br&gt;
| | (darija_buddy.py) | | (gemma3:4b Model) | |&lt;br&gt;
| +-------------------+ +--------------------+ |&lt;br&gt;
+-----------------------------------------------------------+&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The Core Model: Google Gemma 3 (4B)
&lt;/h3&gt;

&lt;p&gt;I selected &lt;strong&gt;Gemma 3 (4B)&lt;/strong&gt; as the reasoning engine. For a compact 4B-parameter open-weight model, Gemma 3 showcases remarkable multilingual understanding, easily deciphering transliterated Arabizi and grasping cultural Moroccan idioms.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Local Inference with Ollama
&lt;/h3&gt;

&lt;p&gt;Instead of relying on remote APIs with variable latency and pay-per-token pricing, the model runs via &lt;strong&gt;Ollama&lt;/strong&gt;. The quantized 4B weights run smoothly on consumer hardware (CPU-only), keeping RAM consumption under ~4 GB.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Gradio Interface
&lt;/h3&gt;

&lt;p&gt;The frontend is encapsulated in a single Python script using &lt;code&gt;gradio.ChatInterface&lt;/code&gt;, making it instantly accessible in any local browser window at &lt;code&gt;http://127.0.0.1:7860&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Pedagogical System Prompt Engineering
&lt;/h3&gt;

&lt;p&gt;The behavior is governed by a system prompt (full prompt in darija_buddy.py on GitHub) enforcing: authentic Darija instead of MSA, short French explanations, Arabizi/French/English input handling, gentle error correction, and short conversational replies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo &amp;amp; Interaction
&lt;/h2&gt;

&lt;p&gt;Here is a practice session where the user initiates in Arabizi, receives natural conversational responses, and gets immediate French vocabulary breakdowns:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr0w4ojmo3kg0mhu3am2i.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr0w4ojmo3kg0mhu3am2i.jpeg" alt=" " width="800" height="248"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhe6b82ahcg2nwuoc0my6.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhe6b82ahcg2nwuoc0my6.jpeg" alt=" " width="800" height="248"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fop1ynjn3m0c5tfu7kxkb.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fop1ynjn3m0c5tfu7kxkb.jpeg" alt=" " width="800" height="248"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Why Open Innovation Matters for This Project&lt;br&gt;
This project highlights why open-weight models and local inference triumph over proprietary closed APIs:&lt;/p&gt;

&lt;p&gt;Zero Operational Cost: Language practice requires repetitive, daily micro-conversations. Running Gemma 3 locally means zero token costs, no API credits expiring, and no credit card requirements for students.&lt;br&gt;
Total Privacy for the Learner: Practicing a new language involves vulnerability and personal conversations. All inference happens in-memory on the laptop; no conversation logs or personal data are ever uploaded to remote commercial servers.&lt;br&gt;
Offline Reliability: University Wi-Fi and mobile data can be unpredictable. Darija Buddy works entirely on a plane, on a train, or in a cafe without an active internet connection.&lt;br&gt;
Customizability: With open weights, I'm not locked into proprietary censorship or forced model upgrades. I can easily fine-tune Gemma on dialectal corpora or swap weights as newer open models release.&lt;br&gt;
What My Friend Said&lt;br&gt;
When I showed it to Alex on a video call:&lt;/p&gt;

&lt;p&gt;"C'est magnifique... et ça va beaucoup m'aider à apprendre le Darija."&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Best Use of Gemma&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Code Repository&lt;br&gt;
The code is completely open-source and easy to reproduce:&lt;/p&gt;

&lt;p&gt;GitHub Repository: &lt;a href="https://github.com/SoufianeZaari/darija-buddy" rel="noopener noreferrer"&gt;https://github.com/SoufianeZaari/darija-buddy&lt;/a&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>DarijaBench: Do AI Models Actually Understand Moroccan Darija?</title>
      <dc:creator>Soufiane Zaari</dc:creator>
      <pubDate>Thu, 01 Oct 2026 16:39:56 +0000</pubDate>
      <link>https://dev.to/soufianezaari/darijabench-do-ai-models-actually-understand-moroccan-darija-4cm1</link>
      <guid>https://dev.to/soufianezaari/darijabench-do-ai-models-actually-understand-moroccan-darija-4cm1</guid>
      <description>&lt;p&gt;Update (Oct 2): Following reviewer feedback, the 0.95 vs 0.90 gap is not statistically significant (exact McNemar p = 0.25) — the honest reading is a four-way tie at n = 60.&lt;/p&gt;

&lt;p&gt;DarijaBench: Do AI Models Actually Understand Moroccan Darija?&lt;br&gt;
As a Moroccan student, I use AI assistants every day. They are brilliant in English and French — but I kept noticing something: ask them something in Darija (Moroccan Arabic dialect, spoken by 35+ million people), and the confident answers start wobbling. So I decided to stop guessing and start measuring. I built DarijaBench, a 60-item benchmark that tests whether today's frontier models truly understand Moroccan Darija — and ran it on four of them.&lt;br&gt;
🔗 Benchmark: &lt;a href="https://www.kaggle.com/benchmarks/soufianzaari/darijabench" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/soufianzaari/darijabench&lt;/a&gt;&lt;br&gt;
What DarijaBench tests&lt;br&gt;
Darija is a low-resource dialect: it is barely present in training data compared to English, French, or even Modern Standard Arabic. DarijaBench probes three practical skills, 20 items each:&lt;br&gt;
Darija → French translation — can the model convert everyday Moroccan sentences into French?&lt;br&gt;
Sentiment analysis — can it tell whether a Darija text is positive, negative, or neutral?&lt;br&gt;
Moroccan cultural knowledge — can it answer questions asked in Darija about Morocco (the capital, couscous ingredients, the 2030 World Cup hosts...)?&lt;br&gt;
Each item is graded automatically with keyword checks, and the final score is the average of the three categories. The whole thing was built with the Kaggle Benchmarks SDK, entirely on the free quota.&lt;br&gt;
The results&lt;br&gt;
Model&lt;br&gt;
DarijaBench score&lt;br&gt;
Gemini 3.7 Flash&lt;br&gt;
0.95&lt;br&gt;
Claude Sonnet 5&lt;br&gt;
0.95&lt;br&gt;
Claude Opus 4.7&lt;br&gt;
0.95&lt;br&gt;
GPT-5.6 Luna&lt;br&gt;
0.90&lt;br&gt;
What I learned&lt;/p&gt;

&lt;p&gt;The top models have mostly cracked basic Darija. A three-way tie at 0.95 was not what I expected — I assumed Darija would still be a weak spot. Frontier labs are clearly doing better on dialectal Arabic than their reputation suggests.&lt;br&gt;
Sentiment is the easiest skill. Gemini scored a perfect 20/20 on sentiment. Emotional tone ("هاد المطعم رائع!" vs "الخدمة خايبة") seems to transfer across dialects and languages with little friction.&lt;br&gt;
Translation is the hardest. At 85%, Darija→French was the lowest category for Gemini. Darija is full of French loanwords used differently than in French, plus idioms with no literal equivalent — exactly where word-level understanding breaks down.&lt;br&gt;
Small gaps need bigger samples. GPT-5.6 Luna's 0.90 vs 0.95 is ~3 more failures on 60 items — but a paired test shows this isn't statistically significant (p = 0.25). Detecting a 5-point gap reliably needs ~200–300 items. That's the main lesson for v2.&lt;br&gt;
Honest caveats. Keyword-based grading is brittle — a correct answer phrased unexpectedly can fail. Sixty items is small. And the three-way tie tells me the benchmark is currently too easy to separate the best models. The next version needs harder items: idioms, heavy code-switching (Darija/French/Spanish mixes are everywhere in the north), and regional variants. Try it yourself The benchmark is public — run more models on it, or fork the task and make it harder. If you speak a low-resource dialect, consider building your own version: the Kaggle Benchmarks SDK makes it surprisingly painless, and measuring is always better than vibing.&lt;br&gt;
Built for the DEV x Kaggle Benchmarking Challenge. &lt;/p&gt;

&lt;h1&gt;
  
  
  kagglechallenge
&lt;/h1&gt;

</description>
      <category>kagglechallenge</category>
      <category>llm</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
