This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built this for my mother, who keeps the household accounts in a paper diary and forwards me bank SMS with the same question: "ye kya tha?" Was this the gas cylinder or the electricity bill? Half the diary entries say "online" and nothing more.
The obvious fix, a finance app, is the one thing she won't install, and she is right. Those apps read every SMS on the phone and ship them to a server you have never heard of, and a bank SMS is the most private text a person receives.
Kharcha (Hindi for "expense") is a small ledger built to run entirely on a laptop. You paste a bank or UPI SMS, or you type what you would have said out loud ("aaj sabzi wale ko 80 diye"), and a 4-billion-parameter open-weight model turns it into one clean row: who, how much, which account, which date, which category. It learns from corrections. Change "Sharma Kirana" from shopping to groceries once and it stays groceries. Tell it Rahul is family and transfers to Rahul stop landing in "other".
It understands the formats of HDFC, SBI, ICICI, Axis, Kotak and PNB, the PhonePe, Google Pay and Paytm notifications, and Hinglish with spoken numbers like "dhai hazaar" and "baarah sau".
Demo
Hosted demo: https://kharcha-4aax.onrender.com (free tier, first load takes about a minute to wake up, and the ledger resets on every deploy). The demo serves the same fine-tuned adapter from Tinker's sampling API because the free tier cannot hold a 4B model.
In the recording: a bank SMS and two spoken notes become ledger rows, "Sharma Kirana" gets corrected from food delivery to groceries once and the app says yaad rakh liya (remembered), and a Hinglish question gets a Hinglish answer.
Code
Kharcha · ghar ka hisaab, phone se bahar nahi jaata
A tiny, private expense ledger for my parent. Paste any Indian bank/UPI SMS, or say "aaj sabzi wale ko 80 diye", and a 4B open-weight model fine-tuned on Tinker turns it into a categorised ledger row The model runs offline on a laptop with Ollama; bank SMS never leave the machine Backboard remembers the corrections ("Sharma Kirana is groceries", "Rahul is my son") and answers Hinglish questions about the month. The demo is hosted on Render: https://kharcha-4aax.onrender.com (free tier, ~1 min cold start).
Built for the DEV Hacktoberfest Weekend Challenge "Build for a Friend" (Oct 2026).
How it works
SMS / voice note ─▶ Qwen3.5-4B + Kharcha LoRA ─▶ {"direction","amount","counterparty","channel","account_last4","date","category"}
▲ hints │
Backboard memory ◀── corrections ───────┘
-
data/– program-labelled dataset generator (real HDFC/SBI/ICICI/Axis/Kotak/PNB/PhonePe/GPay/Paytm formats + Hinglish voice notes). Labels are exact because the program that writes…
How I Built It
The core is a LoRA fine-tune of Qwen3.5-4B trained on Tinker and exported as a 146 MB adapter that I own. Around it: a FastAPI app with a single HTML page, Backboard as the memory layer, and Render hosting the demo.
1. A dataset with no labelling errors
Hand-labelling 1,500 SMS in a weekend was not going to happen, and asking a big model to label them would have baked its mistakes into mine. So I wrote the SMS the way the banks write them. Twenty-six templates copied from real formats, filled with real merchant names, Indian number grouping (1,72,300.00), six date styles, VPAs, reference numbers, and Hinglish voice notes with word numbers. The program that writes the message also writes the label, so the ground truth is exact by construction. A quarter of the training messages are then corrupted the way phones corrupt them: truncated, lower-cased, punctuation stripped.
About one in eight examples carries a "Known rules from the user" block in the system prompt, such as Treat Agarwal Sabzi Bhandar as groceries. or Rahul Kumar is family (beta). That is how the model learns to obey the memory layer rather than its own prior.
1200 train / 150 val / 150 test · 12 categories · 26 message templates
2. Fine-tuning on Tinker
Tinker gives you a training loop, not a black box. The whole run is forward_backward plus optim_step on batches of rendered conversations:
training_client = service.create_lora_training_client(base_model="Qwen/Qwen3.5-4B", rank=16)
renderer = get_renderer("qwen3_5_disable_thinking", training_client.get_tokenizer())
data = [conversation_to_datum(conv, renderer, max_length=640,
train_on_what=TrainOnWhat.LAST_ASSISTANT_MESSAGE) for conv in rows]
for step in range(111):
fb = training_client.forward_backward(batch(step), "cross_entropy")
training_client.optim_step(adam_params=tinker.AdamParams(learning_rate=lr(step)))
sampler_path = training_client.save_weights_for_sampler(name="kharcha-v1").result().path
Three epochs, 111 steps, rank 16, about 1.1 million training tokens. The run cost under a dollar. The model only sees loss on the JSON tokens; the 266-token system prompt and the SMS are context, not targets.
3. Did it actually beat the baseline?
Same 150 held-out messages, same prompt, exact-match on every field. The two baselines are zero-shot: the untouched Qwen3.5-4B and gpt-oss-120b, a model thirty times larger.
| System | valid JSON | amount | counterparty | date | category | all fields | p50 latency | $ / 1k msgs |
|---|---|---|---|---|---|---|---|---|
| Qwen3.5-4B zero-shot | 93% | 89% | 76% | 72% | 71% | 28% | 2.2 s | $0.33 |
| gpt-oss-120b zero-shot | 97% | 97% | 85% | 76% | 81% | 49% | 2.8 s | $0.38 |
| Qwen3.5-4B + Kharcha LoRA | 100% | 100% | 100% | 100% | 100% | 97% | 2.7 s | $0.31 |
The four misses are all the same thing: whether "jooti ke 1200 lage" was paid in cash or by an unknown channel. My own labels decide that arbitrarily, so I count those as label noise rather than model error.
Two honest caveats. First, part of the gap is convention: the big model writes "ATM" where my label says "SBI Bank ATM LAJPAT NAGAR", and both are defensible. Category accuracy, which has no convention problem, still moves from 81% to 100%. Second, the latency on Tinker is similar across models because it is dominated by queueing, so it says little about the size difference; the number that would matter is the offline one on a laptop, which I did not get to measure.
Where the gap is real: the big model calls a petrol pump "other", a Jio recharge "other", a Croma autopay "subscriptions", and misses the date inside a PhonePe transaction ID. The fine-tuned 4B gets every one, because it has seen a thousand of them.
On the 7 test messages that carried a user rule, the baselines followed the rule 86% of the time. The fine-tune followed it 100%. That is the number that makes the memory layer trustworthy.
4. Bringing the weights home
adapter_dir = weights.download(tinker_path=sampler_path, output_dir="out/adapter_raw")
weights.build_lora_adapter(base_model="Qwen/Qwen3.5-4B", adapter_path=adapter_dir, output_path="out/peft_adapter")
That is a 146 MB adapter_model.safetensors sitting in my out/ folder: the whole of what the model learned, in a file I own. train/export_adapter.py carries it the rest of the way, merging into the base, converting to GGUF and registering it with Ollama, and app/parser.py already has the OllamaBackend that the UI switches to with PARSER_BACKEND=ollama. I ran out of weekend (and of home bandwidth: the base weights are 8 GB) before I could time that last step on the laptop, so the table above shows Tinker-served numbers only. The hosted demo and the laptop run the same code and the same adapter; only that one environment variable differs.
5. Memory with Backboard
Backboard stores facts at the assistant level, so one "household" assistant remembers across threads, devices and sessions. When a category gets fixed in the UI, the app writes one memory:
POST /assistants/{id}/memories {"content": "Treat SHARMA KIRANA STORE as groceries."}
Before parsing the next SMS it runs a semantic search over those memories with the raw message as the query, keeps only hits whose subject actually appears in the text, and prepends them to the system prompt as the Known rules block the model was trained on. Memory writes and reads are plain API calls with no LLM in the loop, which is what you want for something as deterministic as "this merchant is groceries".
The Ask box ("is mahine kirane pe kitna gaya?") pulls the relevant memories plus the month's ledger summary and hands them to an open-weight model: a stock Qwen3.5 in Ollama at home, Qwen3.5-9B through Tinker's sampling API on the hosted demo (it shares the parser's tokenizer, which is what keeps the service inside Render's 512 MB). Backboard's own chat endpoint is wired in too and takes over automatically when the account has LLM credits (mine only has memory credits), but I liked that the fallback keeps the entire stack open-weight. One detail I got wrong first: that chat call ran with memory on, and Backboard dutifully extracted "user received a salary of..." from the ledger summary into long-term memory. It now runs read-only. The ledger is context, never a memory.
6. Render
render.yaml describes one Python web service. Tinker key, sampler path and Backboard key are secrets; PARSER_BACKEND=tinker tells the app to serve the adapter from Tinker's sampling API instead of a local Ollama. The same code, one environment variable apart, runs privately on a laptop or publicly on Render.
Why Does Open Innovation Matter?
The data never has to leave. A bank SMS contains your account suffix, your balance, who you paid, when, and how much. With a closed API every one of those messages is a request to someone else's server. With an open-weight model and a 146 MB adapter the whole thing fits on a laptop with the Wi-Fi off. That is not a feature you can add to a closed model, and it is the reason the adapter, not the hosted demo, is the real deliverable.
Fine-tuning is the product. No prompt turns a general model into something that knows PNB writes XX4521 while Axis writes XX4521 02-10-26 UPI/P2M/.... Three epochs of LoRA did. The run cost less than a dollar, the adapter is mine, and if Qwen3.6-4B comes out next month I change one string and retrain over lunch.
Small beats big when the task is narrow. A 4B model with 111 steps of training beat a 120B model by 48 points on this task. In practice that means a model that fits in 3 GB of RAM on an old laptop instead of a GPU cluster.
Memory you can read. Every rule the app learns is a sentence in Backboard that you can list and delete. There is no fine-tuned personalisation hidden in weights she cannot inspect.
Where a closed model would have been better: the Ask box. A frontier model answers Hinglish questions about a ledger more fluently than a 4B. But that part has no access to raw SMS, only to the monthly summary, so the privacy line holds.
My Agent Session
Built over one weekend with Claude Code: the dataset generator, the Tinker training and eval scripts, the app, and most of this post's numbers came out of that session.
The hand-over
We went through one diary page together. The gas cylinder came out as utilities and was immediately disputed: that is rasoi, kitchen. One click, "Yaad rakh liya", and the next gas SMS landed in groceries on its own.
The verdict: "Theek hai. Par diary bhi rakhungi." Fine. But the diary stays.
I will take that.
Prize Categories
Thinking Machines (Tinker), Render, Backboard.


Top comments (7)
Very cool
Yo.
It works with presets but fails for some reason here:
Wanted to joke around but couldn't do it 😭😭😂
Impressive write up @itskumaryash, hoping for the application to become a daily driver for your mom.
Couple of quick questions
what kind of mobile devices can handle this?
How much of knowledge regarding AI agents or Gen AI in general is required before picking up problem statements like these?
The read-only change for Ask is a useful boundary. I'd check the memory lookup separately before calling the laptop path offline: the post says it sends the raw SMS as a Backboard semantic-search query, even when parsing uses Ollama. A fixture with Wi-Fi disabled and a previously corrected merchant could check whether parsing and remembered categories both work without that API. If Backboard is still remote in that mode, an explicit "local parser, cloud memory" label would make the data path clearer. Is there a local memory backend for the fully offline run?
Amazing!!
Nice!
Hey, have you received the badge yet?