DEV Community

Cover image for Shalom Home: my grandpa's phone launcher, rebuilt to stop annoying him.
ROCI
ROCI

Posted on

Shalom Home: my grandpa's phone launcher, rebuilt to stop annoying him.

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My grandfather Shalom calls me a lot. Usually it is not to chat. It is because his phone moved something.

His WhatsApp icon "jumped" to another page. The shortcut he uses to call my mom, his daughter, disappeared. He did nothing wrong: a long press that lasted a second too long, a swipe that started on an icon instead of the background. Android home screens are built for people who know that icons can be dragged, deleted and stacked into folders. Shalom does not know that, and he should not have to.

So I built him Shalom Home: an Android launcher (the app that is your home screen) built on three ideas.

  1. Nothing moves. One page of big tiles: WhatsApp, Photos, the people he calls. No dragging, no deleting, no folders, no accidental pages. Every change happens in a setup screen hidden behind five taps on the clock and a PIN that only I know. Three wrong PINs lock it.
  2. He can just ask. A big button at the bottom says "דבר איתי" ("talk to me"). He says "תתקשר לבת שלי" ("call my daughter"), or in French, "Appelle ma fille", and the phone does it.
  3. It never surprises him. Every voice action shows a 3-second countdown with a giant red "ביטול" (cancel) button and says out loud what it is about to do.

Around that core, the things he is used to still work, just without the chaos:

  • Up to 30 people he can call by voice, each with the words he uses for them ("הבת שלי", "ma fille", a nickname).
  • His own wallpaper shows behind the tiles, so the grandkids' photo stays.
  • A second page for the Contacts "Direct dial" shortcuts and widgets he relies on. I add them from a picker with real previews. If any app tries to put something on his screen, it waits for my approval.
  • Setup in Hebrew or English, so anyone in the family can manage it.

Demo

A 36-second walkthrough: the locked home screen, a Hebrew voice request with its countdown, the second page, and the PIN-protected setup. The APK is on the GitHub release page if you want to try it on an Android phone.

Home, on his wallpaper Second page Adding a widget Setup, in Hebrew
Home Second page Widget picker Setup

Results on my test phrases

I generated Hebrew and French test phrases with a neural TTS voice and ran them through the full pipeline, with a test contact called "Michal" (number 000) whom he calls "הבת שלי", my daughter. In test mode the app only logs what it would do; it never dials.

He says Whisper heard Result
תתקשר למיכל (call Michal) תתקשר למי חל Michal, keywords
אני רוצה לדבר עם הבת שלי (I want to talk to my daughter) exact Michal, keywords
תחייג לילדה (dial the girl) exact Michal, Gemma
Appelle ma fille (call my daughter) exact Michal, Gemma
איפה התמונות שלי (where are my photos) איפות מונות שלי Photos, keywords
מה השעה עכשיו (what time is it) מה שעכשיו nothing: Gemma decided it is not a request

On the emulator, keyword requests finish in about a second and Gemma requests in 2 to 6 seconds. On a real Samsung phone it is pretty fast.

Code

GitHub logo RoeeIlouz / shalom-home

A locked, voice-driven Android home screen for my grandpa. Whisper + Gemma run on-device, in Hebrew.








ontology true
type readme
domain saba-home
status active
tags
android
launcher
whisper
llama-cpp
gemma
accessibility
hacktoberfest
summary Locked, voice-driven Android home screen for an elderly grandparent, with on-device Whisper and Gemma

Shalom Home

A home screen for my grandfather Shalom's Android phone.

Download the APK and the demo video (v1.0)

He kept calling me because a shortcut had moved to another page, or because the button that calls my mom had disappeared. Shalom Home replaces his launcher with one page of big tiles he cannot drag, delete or swipe away, plus a "דבר איתי" ("talk to me") button: he says what he wants in Hebrew or French, and the phone does it.

All speech recognition and understanding runs on the phone with open-weight models. His voice and his contacts never leave the device, and nothing needs an internet connection after setup.


















Home Second page Widget picker Setup (Hebrew)
Home Second page Widget picker Setup

How it

…

Source, build instructions and the APK: https://github.com/RoeeIlouz/shalom-home

How I Built It

microphone
  -> Whisper small (whisper.cpp)          speech to Hebrew text, on the phone
  -> fuzzy Hebrew keyword matcher         instant for "call Michal", "open WhatsApp"
  -> Gemma 3 1B (llama.cpp), if unclear   "I want to talk to the girl" -> his daughter
  -> 3 s spoken countdown + cancel        -> place the call / open the app
Enter fullscreen mode Exit fullscreen mode

It is a Kotlin and Jetpack Compose app with two small native libraries: one wraps whisper.cpp, the other llama.cpp. The models are open weights from Hugging Face, downloaded once in setup: Whisper small (190MB) and Gemma 3 1B quantized to 4 bits (800MB). Both run on the phone's CPU.

Whisper does the listening. Two tricks made it usable for short commands:

  • The decoder is primed with the names of his contacts (initial_prompt), so a spoken name comes out spelled the way it is saved.
  • The encoder normally processes a fixed 30-second window. For a two-second command I shrink audio_ctx to the length of the clip. On my test emulator, that took transcription from 25 to 45 seconds down to about one second.

The keyword matcher handles what Hebrew speech recognition gets wrong in predictable ways: prefixes glued to words (ל, ה, ו: "לאמא" is "to mom"), final-letter forms, spelling variants like אמא / אימא, and Whisper occasionally splitting one word in two ("למי חל" for "למיכל"). If the match is clear, the language model never wakes up.

Gemma 3 1B handles everything without a keyword: "dial the girl", "I want to talk to my daughter", "Appelle ma fille". This is the part I am proudest of, because of how it is constrained:

// The model may only answer with the id of one of HIS options, or "none".
val grammar = "root ::= \"none\" | \"call0\" | \"call1\" | \"app0\" | ..."
val answer = llm.chat(system, "He said: \"$transcript\"", grammar, maxTokens = 6)
Enter fullscreen mode Exit fullscreen mode

llama.cpp's grammar sampler makes it impossible for the model to output anything else. It cannot invent a phone number, a contact, or an action. The worst it can do is pick the wrong one of his people, and the countdown catches that.

A 1B model has habits, and I found them by testing. Asked "what time is it?", it confidently picked the first app in the list. The fix was a separate yes/no question first ("is this a request to call someone or open an app?"), also grammar-constrained. Now small talk is ignored instead of acted on.

Things that went wrong

  • My first test dialed a number. My test harness let the countdown finish, and the emulator placed a (simulated) call to the fake contact I had created. Nobody was called, but it was the most important bug of the weekend: test mode now stops before any action, and only a real tap or a confirmed voice request can dial.
  • Two native libraries, one process. whisper.cpp and llama.cpp each ship their own copy of ggml. I build each into its own JNI library with ggml linked statically and its symbols hidden (-Wl,--exclude-libs,ALL), so they never collide.
  • Samsung is not the emulator. Two things that worked perfectly on the emulator broke on a real Galaxy. The widget picker crashed because I grouped widgets by app name, and several Samsung apps share one; grouping by package fixed it. And Samsung reported no shortcut permission even for the default home app, so "Direct dial" got a fallback that opens the Contacts shortcut screen directly. Testing on the real phone early would have saved me an evening.

Why Does Open Innovation Matter?

I could not have built this on a closed API, and not just because of cost.

  • His voice stays in his pocket. Shalom's requests contain the names of his family. With Whisper and Gemma running on the phone, none of it is sent anywhere. After the one-time model download, the phone does not need the internet at all.
  • It costs nothing to run, forever. He is retired. A voice assistant that needs a subscription, or that stops working when an API changes its pricing, is a voice assistant he will lose.
  • I could constrain the model's output. Grammar-constrained decoding works because I own the sampling loop. That single feature is what makes it safe to let a small model decide who to call.
  • I can swap in a better Hebrew model. The Israeli open-source project ivrit.ai publishes Whisper models fine-tuned on Hebrew speech. Because the app loads any whisper.cpp model file, upgrading Shalom's Hebrew recognition is a file copy, not a vendor negotiation.
  • It runs on his phone. A mid-range Galaxy A with 6GB of RAM runs both models on the CPU. No GPU, no server, no account.

Prize Categories

No partner categories: Shalom Home runs entirely on the phone, so there is nothing to host, no tabular data to forecast and no fine-tuning run. This is an entry for the overall prize.

Disclaimer: Some of the code was generated by Claude Code.

Top comments (0)