This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
In a lot of families there's a pile on the corner of the kitchen table: envelopes that were opened, read twice, and set aside "for when someone can look at it." One of my parents has that pile.
Most of what lands in a pile like that is harmless: a statement, an ad, a thank-you from the health plan. The trouble is that the one that matters looks just like the others. A water shut-off date. Or a letter on very official-looking letterhead saying your Social Security number has been suspended and you can fix it today with Google Play gift cards, as long as you don't tell your bank.
Even in your first language, mail is hard: "Explanation of Benefits — THIS IS NOT A BILL — You may owe $35.00", "past-due balance", "validation notice", "response requested" on a mailer that's actually an ad. In a second language it's harder. The real skill is knowing which of the twelve envelopes is the one that matters.
So this weekend I built Kitchen Table.
What I Built
Kitchen Table is a private letter explainer for someone who gets mail they can't fully read.
They open it on their phone, tap one big button, and photograph the letter. A computer in the same house runs Gemma 4 locally and, a few seconds later, answers in the language they read best:
- The verdict first, in color: Nothing to do. You can put it away. / You need to do something. / Important. Do this soon. / Careful. This looks like a scam.
- What it says, in two to four plain sentences. "Like a kind grown-up child explaining at the kitchen table" is literally in the prompt.
- What to do, by when, with days left and an Add to calendar button.
- The money: how much, and whether they owe it or are getting it back.
- Hard words explained. "Past due", "payment arrangement" and "reconnection fee" each get one plain sentence.
- Warning signs when something is off, and the only advice that matters: Don't pay. Don't call the numbers in the letter. Call Yaro first.
- Questions, typed or spoken. They can hold up the phone and ask, in their own language, "what happens if I don't pay by the 8th?" and hear the answer read back.
- "Send to Yaro": a three-line English summary for me, so our call starts with "I saw the water letter," not "read me the top part again."
The photo is thrown away after reading. The text stays in a SQLite file on the home computer. Nothing goes to a cloud API, there's no account, and it keeps working when the internet is down (only the home Wi-Fi is needed). Setup is one command (./start.sh), which prints a QR code. The parent's phone scans it once, and from then on it's the only device besides the computer that can read their letters. A guest on the Wi-Fi gets nothing.
Demo
An 86-second walkthrough built from real screens of a real run, with English captions (the app itself is in Russian here). You'll hear the actual spoken Russian question the app answered. The narration is Piper, an open-source text-to-speech engine that also runs locally.
The case I actually lose sleep over:
Every letter you see here, and every letter in the test set, is from an organization I invented. Phone numbers are in the reserved 555-01xx range.
Code
pyaroslav
/
kitchen-table
Snap a photo of a confusing letter, get a plain answer in your own language. Local Gemma 4 via Ollama; the mail never leaves the house.
Kitchen Table
Snap a photo of a confusing letter. Get a plain answer in your own language. The letter never leaves the house.
Kitchen Table is a small app for a parent who gets mail they can't fully read: utility bills, insurance statements, jury summonses, "FINAL NOTICE" envelopes, and the scams that are designed to look exactly like them. They take a photo on their phone; a computer at home runs Gemma 4 locally through Ollama and answers:
- a traffic-light verdict: nothing to do, you need to do something, important, do it soon, or careful, this looks like a scam,
- what the letter says, in 2–4 plain sentences in their language (18 languages),
- what to do, by when (with an Add to calendar button), and how much money is involved,
- the hard words explained, the warning signs if it's suspicious,
- answers to follow-up questions, typed or…
Python (FastAPI) on the home computer, a phone page with no framework and no CDN (it has to work offline), SQLite and Ollama. 27 model-free tests run in CI on every push, plus an evaluation harness over 30 synthetic and 9 real letters.
How I Built It
Two passes, and code owns the scary parts
photo ─▶ quality check ─▶ Gemma 4 (vision): exact transcript
scam rules on the transcript
Gemma 4 (JSON schema): explanation in their language
code: date maths, "days left", verdict overrides
Pass 1 only transcribes: "copy names, dates, amounts and phone numbers exactly; write [unreadable] for anything you can't read." Pass 2 never sees pixels. It gets the transcript, today's date and the reader's language, and has to fill a JSON schema: verdict, headline, summary, steps, deadlines, money, glossary, warning signs, and a note for me. Ollama enforces the schema, so the phone never receives half a sentence of markdown.
Splitting the passes gave me three things. Follow-up questions have a transcript to stay grounded in. Each step can be evaluated on its own. And code gets to check the model's work before anyone sees it.
That's the design decision I care most about. The model reads and explains. Code owns everything a parent should never have to trust a model on:
-
Dates. The model returns
2026-10-08. Python works out "6 days left" every time the letter is opened, not once on the day it was read. - Scams. A short list of regex rules runs on the transcript. A strong hit forces the scam warning, whatever the model said: payment by gift card or crypto, "don't tell your bank", arrest threats, "your SSN is suspended", or an official-looking letter whose fine print admits it isn't the government. A suspected scammer's phone number is shown struck through, never as a tap-to-call link.
- Urgency. Anything due within 7 days is Important.
Here's the overriding logic, lightly trimmed. It runs after the model and before the phone:
if redflags.has_strong(hits): # gift cards, crypto, "don't tell your bank", ...
verdict = "scam_warning"
elif verdict in ("file_it", "action_needed") and money_direction != "you_receive":
upcoming = [d for d in deadlines if d["days_left"] >= 0]
if verdict == "file_it" and upcoming and steps:
verdict = "action_needed"
if verdict == "action_needed" and upcoming and upcoming[0]["days_left"] <= 7:
verdict = "urgent"
if verdict == "scam_warning":
deadlines = [] # a scammer's "deadline" is pressure, not an appointment
Why Gemma 4, specifically
I needed one open model family that could do four jobs: read a crooked photo, write fluent plain Russian (or Tagalog, or Persian), return strict JSON, and listen to a spoken question. Gemma 4 does all four, so the whole app is one model family behind one local API:
| gemma4:e4b | gemma4:26b | |
|---|---|---|
| Size on disk (Ollama, Q4_K_M) | 6.6 GB, 7.5B params | 18 GB, 25.2B params (mixture-of-experts, 8 of 128 experts active) |
| Sees photos | ✓ | ✓ |
| Hears audio | ✓ | ✗ |
| Role in Kitchen Table | default; laptop-friendly; the "ears" | best reader if you have the GPU |
| Time per letter (RTX 5090) | ~2.7 s | ~3.5 s |
| Time per letter (CPU only, Ryzen 9 9950X) | ~39 s | not tested |
The mixture-of-experts design is why the 25B model is only about a second slower than the 7.5B one. Both fit on the same GPU at the same time, which is what lets 26b read while e4b listens.
The 2020 bug
I wanted numbers, not vibes. So I wrote 15 letters, rendered them with headless Chrome, and "photographed" them in code: a wood-grain table, a few degrees of rotation, perspective skew, a shadow across one corner, JPEG compression. Both Gemma 4 sizes I was planning to use (e4b and 26b) got all 15 right. That's how you know your test is too easy.
So I wrote nine nastier ones:
- a deadline that only appears on page two;
- a coverage-ending date buried in the fine print of a cheerful newsletter;
- a Spanish-language school letter;
- a photo taken sideways;
- one taken in bad evening light with a shaky hand.
On that last photo, gemma4:e4b read the year 2026 as 2020 and $84.17 as $64.17. Then it explained, confidently and in fluent Russian: "This is an old water bill and doesn't require any action."
That's the worst failure this app can have. It's wrong and calm.
Better prompting couldn't have caught that. A plain photo check could: brightness, contrast and edge sharpness flagged that photo and none of the other 30. A flagged photo gets auto-contrast and sharpening, which fixed the amount but not the year. More importantly, a bad photo can never produce "Nothing to do" any more. It produces "I'm not sure. The photo is too dark or blurry. Please take another one in good light." The bigger gemma4:26b read the same dark photo correctly. The guard stays anyway.
Measuring it
30 synthetic letters in three sets. I wrote the last six after I'd finished changing prompts and rules, ran them once, and didn't touch anything afterwards.
| Model | Right verdict | Never unsafe | Scams caught | Deadlines found | Answer in the right language | Time / letter |
|---|---|---|---|---|---|---|
| gemma4:26b | 30/30 | 30/30 | 8/8 | 100% | 100% | ~3.5 s (GPU) |
| gemma4:e4b | 29/30 | 30/30 | 8/8 | 92% | 100% | ~2.7 s (GPU), ~39 s (CPU only) |
| gemma4:e2b | 25/30 | 25/30 | 6/8 | 88% | 3% | ~2 s (GPU) |
Neither e4b nor 26b raised a single false scam alarm on the 21 legitimate letters, including a debt collector's validation notice and a real water-utility "FINAL NOTICE", which are exactly the letters that look like scams.
Then, real letters
Synthetic letters only prove so much, so once all the tuning was finished I pulled 9 real documents from the agencies and counties that published them:
- three IRS notices, including a two-page "intent to seize your property" letter;
- a federal jury summons;
- a utility's 10-day shut-off notice;
- two fake "Social Security" letters the SSA Inspector General published as warnings;
- two deceptive "deed records" mailers that county offices warned residents about.
I ran them once and changed nothing afterwards.
gemma4:26b got all 9 right. It returned the exact IRS amounts ($1,075.21; $9,533.53) and pay-by dates, called the levy notice urgent, and caught all four scams.
gemma4:e4b got 8. Its miss is the one I'd least like to have: it transcribed a "COUNTY DEED RECORDS — FINAL RENEWAL NOTICE" home-warranty mailer word for word, and then explained it as a real renewal that needs a $199 response. None of my rules fire on it, because it has no gift card and no government disclaimer. I could add a rule now, but I won't: that would turn my only untuned real-world number into a tuned one. So the recommendation is honest instead. e4b is the floor, and if there's a GPU in the house, run 26b.
Four things I didn't expect:
- The scam rules barely mattered for the models I'd actually use. e4b and 26b flagged every scam on their own. On the tiny e2b, the model alone caught zero of the three hard-set scams, and the rules caught two. That's what a seatbelt is for.
- My rules made a weak model worse once. e2b invented deadlines for a library book-sale flyer, and my "upcoming deadline means act" override turned put it away into Important. With the bigger models the overrides only ever corrected things, but I wouldn't have known without the held-out set.
- One rule exists because of this eval, and I'm telling you so. e4b first called an "$89 certified deed copy" mailer you need to do something. Those mailers are legal, so no money-transfer rule fires. Their tell is the disclaimer they're required to print: "not affiliated with any county or government agency." That's now a rule. The held-out set includes a different look-alike, a homestead-exemption "filing service", and both models handled it.
- e2b won't answer in Russian. Ask it to explain in Russian or Ukrainian and it explains in English. For this user that rules it out, whatever its other scores.
The full method is in EVAL.md, with every caveat (synthetic letters, typeset fonts, no handwriting, my own judgement of what counts as acceptable) and a script that regenerates the table.
Voice, and the model that can't hear
Gemma 4 takes audio in the same field as images, so a spoken question is just a 16 kHz WAV sent with the transcript. The phone records it and converts it to WAV itself, and the recording only ever goes to the home computer.
Then I found that gemma4:26b accepts images but not audio; only the small "edge" models can hear. So Kitchen Table checks what the model can do. If the reading model can't hear, gemma4:e4b transcribes the question and 26b answers it: two open models, each doing the part it's good at, on one GPU, joined by about fifteen lines of Python. I tested it end to end with a Russian question synthesized by Piper, a local text-to-speech engine. Gemma transcribed it word for word.
One honest wrinkle: phone browsers only allow in-page microphone recording over HTTPS. On plain home Wi-Fi, the 🎤 button opens the phone's own voice recorder instead, and the file goes through the same path.
The interface is for them, not for me
- Base font size 20px.
- One primary button per screen; the main buttons are 64px tall and nothing tappable is under 48px.
- Colors that mean the same thing in every language.
- A setup screen that asks "Which language do you read best?", with each language written in itself.
- English, Russian, Spanish and Ukrainian strings are hand-written. The other 14 languages are translated once by Gemma itself and cached, with a check that no
{placeholder}got lost on the way. - Arabic and Persian flip the layout to right-to-left, and dates follow the reader's own calendar.
- Dark mode follows the phone.
Lighthouse (mobile) scores the home and letter screens 100 for accessibility.
Why Does Open Innovation Matter?
Look at what's in that pile on the kitchen table: prescriptions, Medicare numbers, bank balances, a debt collector's letter, a court date. It's a map of an older person's life, and it's exactly the information scammers buy. Here's what the open stack gave me that a hosted API couldn't:
- The mail never leaves the house, and that includes voice. There's no cloud speech-to-text and no voice API: Gemma 4 e4b hears the question on the home computer, and the phone's own text-to-speech reads the answer aloud. Beyond that, it isn't "we don't train on it" or "deleted after 30 days": the letter physically never leaves. The photo's GPS metadata is stripped before even the local model sees it, and the photo is discarded after reading.
- It works when the internet doesn't. A home computer and home Wi-Fi. No API key to expire, no service to sunset, no price change.
- It costs nothing per letter. Nobody has to wonder whether a question is "worth" an API call. They can ask the same thing three times.
- I could look inside and change it. I could see that 26b can't hear and route around it. I could measure that e2b ignores the language instruction and rule it out. I could put the scam decision in code I can read, test and explain, instead of in a policy I can't.
- I can swap models without asking anyone. e4b on an old laptop, 26b on my desktop, whatever Gemma ships next on whatever they have. Same app, one environment variable.
The closed version of this app already exists: "take a photo and paste it into a chatbot." It may even be smarter. But I'd have to teach my parent to upload their Medicare letter to a website. Teaching them the opposite is the whole point.
What it can't do (yet)
- It's tested on 30 synthetic and 9 real letters, all photographed in software. Handwriting, creased paper and window envelopes aren't in the test set.
- On the small model, a deceptive mailer without any classic scam markers can still slip through (see above).
- It doesn't replace me, by design: every unclear or suspicious letter ends with "call Yaro first."
- On a laptop without a GPU it takes the better part of a minute per letter. That's fine for mail, less fine for impatience.
Prize Categories
-
Best Use of Gemma. Gemma 4 runs locally through Ollama and does all the reading, explaining, translating (including the app's own interface) and listening:
gemma4:26bfor vision,gemma4:e4bfor audio, with e2b, e4b and 26b measured head to head on 30 synthetic letters, and e4b and 26b on 9 real ones.
Thanks for reading. If there's a pile like this on a kitchen table in your life, setup is ollama pull gemma4:e4b and one Python command.
— Yaro


Top comments (0)