This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
A fully offline C++ tutor that runs on a phone, built for [friend's name], who is learning C++ from scratch.
Most AI tools hand over the answer, and my friend ended up copy-pasting code without understanding it. Byte-Gemma is a fine-tuned Gemma 3 1B that teaches Socratically. It stays under ~100 words, praises one thing you got right, and asks one guiding question. It runs locally on Android, so it works without internet or an API bill.
Slash commands control how much help you get:
| Command | Behaviour |
|---|---|
/explain |
Concept via analogy, tiny example, one question |
/debug |
Only the first hint |
/hint |
A more specific hint, no full solution |
/solution |
Full fix, explanation, practice question |
/review |
One thing done right, the key issue, you fix it |
/quiz |
One question, waits for your answer |
Demo
Code
- Model: https://huggingface.co/99Amit99/byte_v2-gemma-3-1b-it-v2-GGUF/tree/main
- Repo: https://github.com/amitsinha-cell/byte_v2.git
How I Built It
- Base: Gemma 3 1B (open weights), fine-tuned and quantized to Q4_K_M GGUF.
- Data: [~300] examples generated with Gemma 4 as a teacher via the Gemini API. Validation scripts filtered out low-quality outputs, and I merged in cleaned examples from my friend's real coding mistakes.
- Runtime: llama.cpp-based, with PocketPal/Termux on Android.
- A pivot: My first version was a gamified "Byte" persona. Testing showed it wasn't teaching well, so I rebuilt it around Socratic tutoring.
Evaluation (and what failed)
My friend and I compared Byte-Gemma against stock Gemma 3 1B on struct vs class, pointers vs references, and multithreading. My first attempt gave only Byte-Gemma the tutoring system prompt, which made the result meaningless. So I re-ran it with the identical system prompt, runtime, and questions for both models.
Where Byte-Gemma won: following the tutoring rules.
- 6 of 6 replies stayed under 100 words (about 66 words on average for first replies). The stock model went over in 3 of 3 (131 to 214 words).
- Byte-Gemma asked exactly one question per reply. Stock Gemma asked three in the struct/class test and ended others with "Does that make sense?".
- Speed per token was about the same (about 6.4 vs 6.7 tok/s), but replies about 2.7x shorter mean much shorter waits on a phone.
Where it fell short: accuracy.
- Every Byte-Gemma reply had at least one error, and at least 4 of its 6 code snippets were broken (assigning
"Red"to anint,* = 12;,Car<Car>, and a "multithreading" example that never creates a thread). - It never stated the real struct/class difference (default access), and it wrongly said references can't access an object directly.
- With one-word replies like "classes" or "structure", it said "You are right" and then invented that objects are created "using the
<iostream>header". A tutor that agrees with a confused student and adds a wrong fact is the failure I most want to avoid. - Stock Gemma had errors too (calling a reference "not an alias", a class with an accidentally private constructor), but its multithreading explanation was sound and its pointer example was correct.
Subjective scores: Byte-Gemma 5.5/10, stock Gemma 4/10. Byte-Gemma's lead is smaller than my first run suggested, and it comes from teaching format, not knowledge. Caveats: one run per prompt, 6 Byte-Gemma replies vs 3 stock replies, no follow-up turns for stock Gemma, and code judged by reading rather than compiling.
Neither model is reliable enough to teach unsupervised. Next steps:
- Compile every code block in the training data and reject anything that doesn't build.
- Add multi-turn examples where the student answers with one word and the tutor checks it instead of agreeing.
- Re-run the comparison with fixed temperature and seed, including follow-ups for both models.
Why Does Open Innovation Matter?
- Offline: The tutor runs on a phone with no connection.
- Customizable: I could change the model's behavior through fine-tuning, not just prompting.
- Free: No per-token cost for a student project.
- Transparent: My friend could audit the model, find its errors, and show me where it was wrong. Testing against the base model with the same prompt is also what showed me the fine-tune changed its behavior more than its knowledge.
Prize Categories
Google Gemma


Top comments (0)