This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My parents have never lost money to a scam. I'd like to keep it that way.
My mom lives on WhatsApp. My dad pays for everything with UPI. Their phones get a steady drip of messages written to make someone panic: a KYC deadline "today", a parcel "held at customs", the electricity "disconnected tonight", a "police officer" on a video call. They've been careful so far. But these messages are designed for the one moment you're tired, alone or scared. And they only need to work once.
So I built them Rakshak ("protector"): a second opinion that lives on a laptop in our house. You hand it a suspicious message, a screenshot, or a description of a phone call, and within seconds a calm voice tells you in plain English whether it's real and what to do.
- Mom (Android) shares a message to it straight from her share sheet. A big card comes up: red with a raised hand for This is a scam, amber for Be careful, green for This looks like a normal message, grey for Couldn't check right now. The explanation is read aloud. For a suspicious call, she holds a button and describes it.
- Dad (iPhone) uses two Shortcuts: "Rakshak" in the share sheet, and "Rakshak Bolo" (bolo means "speak") to describe a call out loud. Same card, same voice.
- Me: on a scam verdict my phone gets a push with which parent, what kind of scam, and what time. Never the message.
Every explanation is three short sentences at most: the verdict, the one or two strongest warning signs quoted exactly from the message, and one thing to do. Don't tap. Don't pay. Never share the OTP. Call me.
Nothing leaves the house. That's not a slogan, it's the architecture.
Demo
There's no public link, and that's on purpose. Rakshak runs on a MacBook at home and is reachable only over our private Tailscale network, because the whole point is that my parents' messages never touch someone else's server. Here it is running for real:
Left to right: the home screen; the check in progress; an app-install scam caught with the dangerous link quoted, plus Listen again for the spoken explanation; a real bank OTP correctly left alone.
The hub's status page for me: three local models, latency against targets, today's verdicts per parent, and a "was this right?" queue. Metadata only. It never shows a message.
Code
Rakshak
A scam checker for my parents that runs on a MacBook at home.
My mom uses an Android phone and my dad uses an iPhone. Both get a steady stream of messages written to make a parent panic: fake KYC deadlines, parcels "held at customs", electricity cut off "tonight", fake police calls. Rakshak gives them a second opinion. Mom shares a message or screenshot to the Rakshak web app from her share sheet; Dad shares it to an iOS Shortcut. A few seconds later they see a large verdict card (scam, be careful, looks normal, or couldn't check) and hear a short spoken explanation in plain English. If it's a scam, my phone gets a notification that says which parent and what kind of scam, never the message itself. Every model runs on the Mac. The messages never leave the house.
The verdict comes from a small model fine-tunedβ¦
Everything is in the repo: the rules engine, the LangGraph pipeline, training, export and evaluation scripts, the web app, and the privacy tests. The dataset is regenerated locally and the model weights are built by the export script rather than checked in, but nothing about them is secret. They're trained only on synthetic and public data.
How I Built It
A MacBook Pro (M3 Pro, 18 GB) is the hub. It binds to localhost only; phones reach it over a private tailnet. Each check runs through a small LangGraph graph: normalize, perceive the image or audio, rules, detect, fuse, explain, then notify and store. SQLite keeps the records. Every model runs locally through llama.cpp and Ollama.
Two open models, three jobs
I wanted the verdict to come from a model I'd fine-tuned and measured, not a prompt. Tinker doesn't offer Gemma, which made the split natural. It also gave Gemma a second, less obvious job.
Gemma 4 E2B wrote the training set. I had no scam dataset, and I refused to build one from my parents' private messages. So Gemma generated it: from about 60 short seed patterns across every scam and genuine category, it produced around 1,300 realistic messages in English, Hinglish and Devanagari, varied across channel and the tricks scammers use: misspellings, odd spacing, look-alike characters, shortened links. Every quoted red flag had to appear exactly in the generated text or the item was dropped.
Tinker trained the judge. Rakshak-Detector is Qwen3.5-4B with a LoRA adapter trained on Tinker on that data, then merged, converted to GGUF and quantized to Q4_K_M (2.9 GB) for llama-server. It returns only JSON: verdict, category, red-flag quotes. I read the scam probability from the log-probabilities of the verdict token, not from the text, so the fusion step has a real number to work with.
Gemma explains the verdict. It transcribes screenshots and voice notes on-device, and turns the verdict into a short explanation in plain words. It never decides. It receives the verdict as input, and a post-check rejects any explanation that drops the verdict word, negates it, adds a link or number that wasn't in the message, or runs long. On failure a fixed template takes over.
Under all of it sits a rules engine of plain, unit-tested functions. An .apk link, a look-alike of a bank's domain, or "enter your UPI PIN to receive money" is a scam whatever the model says. Softer signals push toward Be careful: a bank claiming to text from a plain mobile number, a link shortener, arrest threats, urgency plus a payment ask. The word lists cover English, Hinglish and Devanagari, because that's how scams actually arrive.
Two safety rules live in code, not prompts: red flags shown to my parents must be exact substrings of their message, or they're dropped; and if the detector is down the answer is Couldn't check, never looks normal.
Training
LoRA rank 32, loss on the assistant's JSON only, batch 32, 3 epochs, thinking mode off, learning rate 4.91e-4 from tinker_cookbook's get_lr. In 20% of training examples I blanked the rule signals so the model can't just copy the rules. I ran a smoke run, a learning-rate sweep, and then a retrain, because two independent review passes caught the quiet failure of synthetic data: the first model had learned shortcuts like a link means scam and an unknown sender means scam, which would have meant false alarms on real bank SMS. I fixed the generator to balance those cues and trained again. Total: about $3.80 of the $10 Tinker credit.
Results, on data it never saw
I did not dataset my parents' messages, so there is no family test set. Instead, the evaluation is 100 public fraud reports: scams and genuine messages documented by RBI, PIB, India's cybercrime advisories, news outlets, fact-checkers and court records, each with a source URL and verified verbatim on its page (65 scam, 35 genuine). The model never trains on them. 95% bootstrap CIs; scam recall counts any warning on a real scam.
| System | Scam recall | False alarms on genuine | On alarming-but-genuine | Macro-F1 | Valid JSON |
|---|---|---|---|---|---|
| Rules only | 0.43 | 9% | 11% | 0.05 | n/a |
| Gemma 4 E2B zero-shot | 0.86 | 14% | 22% | 0.38 | 0% |
| Qwen3.5-4B base, few-shot | 1.00 | 43% | 89% | 0.54 | 33% |
| Qwen3.5-4B + Tinker LoRA | 0.92 [0.86β0.98] | 9% | 22% | 0.59 [0.55β0.63] | 100% |
| Full Rakshak (rules + fusion) | 0.94 | 14% | 22% | 0.58 | n/a |
The interesting number isn't the F1. It's the false alarms. The base Qwen model with few-shot prompting catches every scam, but only by crying wolf. It flags 43% of genuine messages and 89% of genuine-but-alarming ones (real OTPs, real debit alerts), and only a third of its replies are even valid JSON. Fine-tuning the same model on this one narrow task cuts false alarms on genuine messages to 9% and makes the output 100% machine-usable, while keeping recall at 0.92, on data it never trained on.
That's the whole project in one row. Not missing scams is easy. Not terrifying your mom about a real OTP is the hard part, and that's exactly what a small, fine-tuned open model fixed.
(Adding the rules on top nudges recall to 0.94 but raises genuine false alarms, because a few real government advisories trip the threat lexicon. I can see that trade-off in the numbers, which is the point of measuring.)
These are small-sample numbers on public data, not a benchmark claim. The synthetic dev set is near-saturated and I don't report it.
Speed on a laptop
All three model servers run together on 18 GB. End to end:
- Genuine (detector only): p50 0.8 s, p95 0.9 s
- Scam (detector + spoken explanation): p50 4.1 s, p95 11.1 s
The detector dominates the scam path; the explainer adds about 1.6 s. The scam path is over my 6-second goal. Shorter red-flag output, constrained JSON on the first attempt, and streaming the verdict before the explanation are next.
Tracing without reading anyone's messages
Sentry gets one transaction per check and one span per graph node, so I can see where the seconds go and which model answered, but never the content. Span attributes come from an allowlist (node, model, token counts, latency, verdict, category, p_scam, rule-signal names, retries, error type), and a scrubbing hook deletes any text, quote, explanation, image or audio field at any depth, including inside exception messages. Request bodies and local variables are never captured.
I don't just trust that. A canary test pushes a fake message with a unique marker through every input path of the real app and fails if the marker appears in any Sentry payload, ntfy request, log line or console output, or if the hub talks to anything but the local model servers and ntfy. I broke the scrubber on purpose three ways; the test caught all three.
Why Does Open Innovation Matter?
Privacy that's literally true. The messages my parents get are about their bank, their pension, their family. Asking a hosted API "is this a scam?" means one more stranger reads them. With open weights, the model comes to the data, and checks keep working with the Mac's internet off.
A small specialist beats a big generalist here. "Is this a scam, and which words prove it?" is narrow and structured. Open weights let me fine-tune a 4B model on exactly that and measure it against its own base: false alarms went from 43% to 9%, valid output from 33% to 100%. You can't do that with a model you can't fine-tune or run.
Open models can make the data, too. Gemma generated the whole training set locally. No scraping, no labeling service, no one's private messages.
Zero marginal cost. Each check costs electricity, so there's no reason to ration them. My parents can share every message that feels off.
The honest caveat. Training wasn't fully local: synthetic and redacted public examples went to Tinker's servers to train the adapter. No family message was ever part of it. At runtime, nothing leaves the house.
My Agent Session
I built Rakshak with Claude Code, working from a written requirements document that fixed the schemas, prompts, thresholds and privacy rules before any code existed. The agent did the heavy lifting in parallel (rules engine, hub pipeline, training and export, web app), and the most valuable thing it did was review its own work. Two independent review passes caught the detector throwing away suspicious probability mass, and the synthetic-data shortcuts that would have caused false alarms on real messages. The judgment calls I kept for myself were the ones about honesty: never letting public reports become a "test set", and insisting the evaluation say plainly what data it actually ran on.
Prize Categories
- Best Use of Gemma. Gemma 4 E2B generated the training set, reads screenshots and voice notes on-device, and writes every warning my parents hear, with a post-check that keeps it from ever changing the verdict.
- Best Use of Tinker. The verdict comes from Qwen3.5-4B fine-tuned with LoRA on Tinker, exported to a quantized GGUF, and measured against its own base model and Gemma on public data it never trained on, where fine-tuning cut genuine false alarms from 43% to 9% and raised valid structured output from 33% to 100%.
- Best Use of Sentry Agent Tracing. One trace per check and one span per LangGraph node, with a canary test proving no message text ever reaches Sentry.





Top comments (0)