<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yash Kumar</title>
    <description>The latest articles on DEV Community by Yash Kumar (@coffeeandcommits).</description>
    <link>https://dev.to/coffeeandcommits</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1193656%2Fae1e6149-c9d9-4e2a-a09d-7882022584fa.png</url>
      <title>DEV Community: Yash Kumar</title>
      <link>https://dev.to/coffeeandcommits</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/coffeeandcommits"/>
    <language>en</language>
    <item>
      <title>Kharcha: a 4B model that reads Indian bank SMS so the money stays on your laptop</title>
      <dc:creator>Yash Kumar</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:37:59 +0000</pubDate>
      <link>https://dev.to/coffeeandcommits/kharcha-a-4b-model-that-reads-indian-bank-sms-so-the-money-stays-on-your-laptop-2n8j</link>
      <guid>https://dev.to/coffeeandcommits/kharcha-a-4b-model-that-reads-indian-bank-sms-so-the-money-stays-on-your-laptop-2n8j</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;I built this for my mother, who keeps the household accounts in a paper diary and forwards me bank SMS with the same question: &lt;em&gt;"ye kya tha?"&lt;/em&gt; Was this the gas cylinder or the electricity bill? Half the diary entries say "online" and nothing more.&lt;/p&gt;

&lt;p&gt;The obvious fix, a finance app, is the one thing she won't install, and she is right. Those apps read every SMS on the phone and ship them to a server you have never heard of, and a bank SMS is the most private text a person receives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Kharcha&lt;/strong&gt; (Hindi for "expense") is a small ledger built to run entirely on a laptop. You paste a bank or UPI SMS, or you type what you would have said out loud ("aaj sabzi wale ko 80 diye"), and a 4-billion-parameter open-weight model turns it into one clean row: who, how much, which account, which date, which category. It learns from corrections. Change "Sharma Kirana" from shopping to groceries once and it stays groceries. Tell it Rahul is family and transfers to Rahul stop landing in "other".&lt;/p&gt;

&lt;p&gt;It understands the formats of HDFC, SBI, ICICI, Axis, Kotak and PNB, the PhonePe, Google Pay and Paytm notifications, and Hinglish with spoken numbers like "dhai hazaar" and "baarah sau".&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2bhrxx4u2r97f35q0crr.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2bhrxx4u2r97f35q0crr.gif" alt="Kharcha demo: paste an SMS, say a voice note, fix a category once, ask in Hinglish" width="760" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Hosted demo: &lt;strong&gt;&lt;a href="https://kharcha-4aax.onrender.com" rel="noopener noreferrer"&gt;https://kharcha-4aax.onrender.com&lt;/a&gt;&lt;/strong&gt; (free tier, first load takes about a minute to wake up, and the ledger resets on every deploy). The demo serves the same fine-tuned adapter from Tinker's sampling API because the free tier cannot hold a 4B model.&lt;/p&gt;

&lt;p&gt;In the recording: a bank SMS and two spoken notes become ledger rows, "Sharma Kirana" gets corrected from food delivery to groceries once and the app says &lt;em&gt;yaad rakh liya&lt;/em&gt; (remembered), and a Hinglish question gets a Hinglish answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/its-kumar-yash" rel="noopener noreferrer"&gt;
        its-kumar-yash
      &lt;/a&gt; / &lt;a href="https://github.com/its-kumar-yash/kharcha" rel="noopener noreferrer"&gt;
        kharcha
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Kharcha · ghar ka hisaab, phone se bahar nahi jaata&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;A tiny, private expense ledger for my parent. Paste any Indian bank/UPI SMS, or say
"aaj sabzi wale ko 80 diye", and a 4B open-weight model fine-tuned on
&lt;a href="https://thinkingmachines.ai/tinker/" rel="nofollow noopener noreferrer"&gt;Tinker&lt;/a&gt; turns it into a categorised ledger row
The model runs &lt;strong&gt;offline on a laptop with Ollama&lt;/strong&gt;; bank SMS never leave the machine
&lt;a href="https://backboard.io" rel="nofollow noopener noreferrer"&gt;Backboard&lt;/a&gt; remembers the corrections ("Sharma Kirana is groceries",
"Rahul is my son") and answers Hinglish questions about the month. The demo is hosted on
&lt;a href="https://render.com" rel="nofollow noopener noreferrer"&gt;Render&lt;/a&gt;: &lt;strong&gt;&lt;a href="https://kharcha-4aax.onrender.com" rel="nofollow noopener noreferrer"&gt;https://kharcha-4aax.onrender.com&lt;/a&gt;&lt;/strong&gt; (free tier, ~1 min cold start).&lt;/p&gt;
&lt;p&gt;Built for the DEV Hacktoberfest Weekend Challenge "Build for a Friend" (Oct 2026).&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;How it works&lt;/h2&gt;

&lt;/div&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;SMS / voice note ─▶ Qwen3.5-4B + Kharcha LoRA ─▶ {"direction","amount","counterparty","channel","account_last4","date","category"}
                          ▲ hints                         │
                 Backboard memory  ◀── corrections ───────┘
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;data/&lt;/code&gt; – program-labelled dataset generator (real HDFC/SBI/ICICI/Axis/Kotak/PNB/PhonePe/GPay/Paytm formats + Hinglish voice notes). Labels are exact because the program that writes…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/its-kumar-yash/kharcha" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The core is a &lt;strong&gt;LoRA fine-tune of Qwen3.5-4B&lt;/strong&gt; trained on &lt;a href="https://thinkingmachines.ai/tinker/" rel="noopener noreferrer"&gt;Tinker&lt;/a&gt; and exported as a 146 MB adapter that I own. Around it: a FastAPI app with a single HTML page, &lt;a href="https://backboard.io" rel="noopener noreferrer"&gt;Backboard&lt;/a&gt; as the memory layer, and &lt;a href="https://render.com" rel="noopener noreferrer"&gt;Render&lt;/a&gt; hosting the demo.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. A dataset with no labelling errors
&lt;/h3&gt;

&lt;p&gt;Hand-labelling 1,500 SMS in a weekend was not going to happen, and asking a big model to label them would have baked its mistakes into mine. So I wrote the SMS the way the banks write them. Twenty-six templates copied from real formats, filled with real merchant names, Indian number grouping (&lt;code&gt;1,72,300.00&lt;/code&gt;), six date styles, VPAs, reference numbers, and Hinglish voice notes with word numbers. &lt;strong&gt;The program that writes the message also writes the label&lt;/strong&gt;, so the ground truth is exact by construction. A quarter of the training messages are then corrupted the way phones corrupt them: truncated, lower-cased, punctuation stripped.&lt;/p&gt;

&lt;p&gt;About one in eight examples carries a &lt;em&gt;"Known rules from the user"&lt;/em&gt; block in the system prompt, such as &lt;code&gt;Treat Agarwal Sabzi Bhandar as groceries.&lt;/code&gt; or &lt;code&gt;Rahul Kumar is family (beta)&lt;/code&gt;. That is how the model learns to obey the memory layer rather than its own prior.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1200 train / 150 val / 150 test · 12 categories · 26 message templates
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Fine-tuning on Tinker
&lt;/h3&gt;

&lt;p&gt;Tinker gives you a training loop, not a black box. The whole run is &lt;code&gt;forward_backward&lt;/code&gt; plus &lt;code&gt;optim_step&lt;/code&gt; on batches of rendered conversations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;training_client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_lora_training_client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen/Qwen3.5-4B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;renderer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_renderer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qwen3_5_disable_thinking&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;training_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_tokenizer&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;conversation_to_datum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;conv&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;renderer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_length&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;640&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                              &lt;span class="n"&gt;train_on_what&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;TrainOnWhat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;LAST_ASSISTANT_MESSAGE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;conv&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;111&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;fb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;training_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;forward_backward&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cross_entropy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;training_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;optim_step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;adam_params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;tinker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AdamParams&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;learning_rate&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;lr&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;
&lt;span class="n"&gt;sampler_path&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;training_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;save_weights_for_sampler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kharcha-v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;result&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three epochs, 111 steps, rank 16, about 1.1 million training tokens. &lt;strong&gt;The run cost under a dollar.&lt;/strong&gt; The model only sees loss on the JSON tokens; the 266-token system prompt and the SMS are context, not targets.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Did it actually beat the baseline?
&lt;/h3&gt;

&lt;p&gt;Same 150 held-out messages, same prompt, exact-match on every field. The two baselines are zero-shot: the untouched Qwen3.5-4B and gpt-oss-120b, a model thirty times larger.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;valid JSON&lt;/th&gt;
&lt;th&gt;amount&lt;/th&gt;
&lt;th&gt;counterparty&lt;/th&gt;
&lt;th&gt;date&lt;/th&gt;
&lt;th&gt;category&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;all fields&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;p50 latency&lt;/th&gt;
&lt;th&gt;$ / 1k msgs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-4B zero-shot&lt;/td&gt;
&lt;td&gt;93%&lt;/td&gt;
&lt;td&gt;89%&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;72%&lt;/td&gt;
&lt;td&gt;71%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;28%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.2 s&lt;/td&gt;
&lt;td&gt;$0.33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-120b zero-shot&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;97%&lt;/td&gt;
&lt;td&gt;85%&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;81%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.8 s&lt;/td&gt;
&lt;td&gt;$0.38&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3.5-4B + Kharcha LoRA&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;97%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.7 s&lt;/td&gt;
&lt;td&gt;$0.31&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnakh8v9yoo38fj6wt8ss.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnakh8v9yoo38fj6wt8ss.png" alt="All-fields and category accuracy on 150 held-out messages" width="799" height="373"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The four misses are all the same thing: whether "jooti ke 1200 lage" was paid in cash or by an unknown channel. My own labels decide that arbitrarily, so I count those as label noise rather than model error.&lt;/p&gt;

&lt;p&gt;Two honest caveats. First, part of the gap is convention: the big model writes &lt;code&gt;"ATM"&lt;/code&gt; where my label says &lt;code&gt;"SBI Bank ATM LAJPAT NAGAR"&lt;/code&gt;, and both are defensible. Category accuracy, which has no convention problem, still moves from 81% to 100%. Second, the latency on Tinker is similar across models because it is dominated by queueing, so it says little about the size difference; the number that would matter is the offline one on a laptop, which I did not get to measure.&lt;/p&gt;

&lt;p&gt;Where the gap is real: the big model calls a petrol pump "other", a Jio recharge "other", a Croma autopay "subscriptions", and misses the date inside a PhonePe transaction ID. The fine-tuned 4B gets every one, because it has seen a thousand of them.&lt;/p&gt;

&lt;p&gt;On the 7 test messages that carried a user rule, the baselines followed the rule 86% of the time. The fine-tune followed it 100%. That is the number that makes the memory layer trustworthy.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Bringing the weights home
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;adapter_dir&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;download&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tinker_path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;sampler_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output_dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out/adapter_raw&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;weights&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;build_lora_adapter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Qwen/Qwen3.5-4B&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;adapter_path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;adapter_dir&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output_path&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out/peft_adapter&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a 146 MB &lt;code&gt;adapter_model.safetensors&lt;/code&gt; sitting in my &lt;code&gt;out/&lt;/code&gt; folder: the whole of what the model learned, in a file I own. &lt;code&gt;train/export_adapter.py&lt;/code&gt; carries it the rest of the way, merging into the base, converting to GGUF and registering it with Ollama, and &lt;code&gt;app/parser.py&lt;/code&gt; already has the &lt;code&gt;OllamaBackend&lt;/code&gt; that the UI switches to with &lt;code&gt;PARSER_BACKEND=ollama&lt;/code&gt;. I ran out of weekend (and of home bandwidth: the base weights are 8 GB) before I could time that last step on the laptop, so the table above shows Tinker-served numbers only. The hosted demo and the laptop run the same code and the same adapter; only that one environment variable differs.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Memory with Backboard
&lt;/h3&gt;

&lt;p&gt;Backboard stores facts at the assistant level, so one "household" assistant remembers across threads, devices and sessions. When a category gets fixed in the UI, the app writes one memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /assistants/{id}/memories   {"content": "Treat SHARMA KIRANA STORE as groceries."}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before parsing the next SMS it runs a semantic search over those memories with the raw message as the query, keeps only hits whose subject actually appears in the text, and prepends them to the system prompt as the &lt;em&gt;Known rules&lt;/em&gt; block the model was trained on. Memory writes and reads are plain API calls with no LLM in the loop, which is what you want for something as deterministic as "this merchant is groceries".&lt;/p&gt;

&lt;p&gt;The Ask box ("is mahine kirane pe kitna gaya?") pulls the relevant memories plus the month's ledger summary and hands them to an open-weight model: a stock Qwen3.5 in Ollama at home, Qwen3.5-9B through Tinker's sampling API on the hosted demo (it shares the parser's tokenizer, which is what keeps the service inside Render's 512 MB). Backboard's own chat endpoint is wired in too and takes over automatically when the account has LLM credits (mine only has memory credits), but I liked that the fallback keeps the entire stack open-weight. One detail I got wrong first: that chat call ran with memory on, and Backboard dutifully extracted "user received a salary of..." from the ledger summary into long-term memory. It now runs read-only. The ledger is context, never a memory.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Render
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;render.yaml&lt;/code&gt; describes one Python web service. Tinker key, sampler path and Backboard key are secrets; &lt;code&gt;PARSER_BACKEND=tinker&lt;/code&gt; tells the app to serve the adapter from Tinker's sampling API instead of a local Ollama. The same code, one environment variable apart, runs privately on a laptop or publicly on Render.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The data never has to leave.&lt;/strong&gt; A bank SMS contains your account suffix, your balance, who you paid, when, and how much. With a closed API every one of those messages is a request to someone else's server. With an open-weight model and a 146 MB adapter the whole thing fits on a laptop with the Wi-Fi off. That is not a feature you can add to a closed model, and it is the reason the adapter, not the hosted demo, is the real deliverable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fine-tuning is the product.&lt;/strong&gt; No prompt turns a general model into something that knows PNB writes &lt;code&gt;XX4521&lt;/code&gt; while Axis writes &lt;code&gt;XX4521 02-10-26 UPI/P2M/...&lt;/code&gt;. Three epochs of LoRA did. The run cost less than a dollar, the adapter is mine, and if Qwen3.6-4B comes out next month I change one string and retrain over lunch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Small beats big when the task is narrow.&lt;/strong&gt; A 4B model with 111 steps of training beat a 120B model by 48 points on this task. In practice that means a model that fits in 3 GB of RAM on an old laptop instead of a GPU cluster.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory you can read.&lt;/strong&gt; Every rule the app learns is a sentence in Backboard that you can list and delete. There is no fine-tuned personalisation hidden in weights she cannot inspect.&lt;/p&gt;

&lt;p&gt;Where a closed model would have been better: the Ask box. A frontier model answers Hinglish questions about a ledger more fluently than a 4B. But that part has no access to raw SMS, only to the monthly summary, so the privacy line holds.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;Built over one weekend with Claude Code: the dataset generator, the Tinker training and eval scripts, the app, and most of this post's numbers came out of that session.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hand-over
&lt;/h2&gt;

&lt;p&gt;We went through one diary page together. The gas cylinder came out as &lt;em&gt;utilities&lt;/em&gt; and was immediately disputed: that is &lt;em&gt;rasoi&lt;/em&gt;, kitchen. One click, &lt;em&gt;"Yaad rakh liya"&lt;/em&gt;, and the next gas SMS landed in groceries on its own.&lt;/p&gt;

&lt;p&gt;The verdict: &lt;em&gt;"Theek hai. Par diary bhi rakhungi."&lt;/em&gt; Fine. But the diary stays.&lt;/p&gt;

&lt;p&gt;I will take that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;Thinking Machines (Tinker), Render, Backboard.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
