DEV Community

Prateek Sharma
Prateek Sharma

Posted on

Stopped pasting grocery list into ChatGPT, Gemma 4 E2B runs it on my GTX 1650

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

We roomies have a group where we put all our grocery expenses. Then someone has to calculate all or we just export the chat to ChatGPT.
Now this project does everything locally using Gemma, no data leaves my machine.

Demo

Code

WAExpenseSplitLocalAI

Split shared grocery spending without pasting your group's money chats into a cloud chatbot Paste a WhatsApp export (and/or freeform lines) β†’ a local open-weight model reads every line β†’ you correct anything in a table β†’ deterministic Python computes who pays whom No API key, no account, no data leaves your laptop.

Built with Gemma 4 E2B (Apache 2.0) running locally via Ollama. Originally made for a Hacktoberfest "Build for a Friend" entry: month-end grocery splits for roommates, parsed on-device for privacy.

Why local

Grocery + money chats reveal food habits, routines, locations, and financial state. Pasting them into an online AI sends all of that to someone else's server. Running the model on your own machine means:

  • Private β€” traffic never leaves localhost:11434.
  • Free β€” zero per-parse cost, unlimited re-runs.
  • Offline β€” works on hostel WiFi or no WiFi after a one-time ~4.3 GB model…

How I Built It

The open core is Gemma 4 E2B (gemma4:e2b-it-qat, 4.3 GB quant) running on my own laptop
through Ollama, with a Streamlit UI and plain Python around it. No API key, no account,
no cloud β€” inference happens on localhost:11434.

The architecture has one rule: code owns structure, the model owns meaning, Python owns money.

  • Code owns structure. WhatsApp export envelopes ([28/09, 21:30] Name: …) are peeled deterministically before the model ever sees them β€” so a timestamp can never be mistaken for an amount. (My first version fed raw lines to the model and it solemnly reported expenses of Rs.28 and Rs.29. Those were the dates.)
  • Model owns meaning. One Ollama call per line β€” I measured that this small model drops lines when batched, so each line gets its own call. It first lists every amount it sees, then returns one item per amount; its own count is checked and retried once on mismatch. Anything still off becomes a visible warning, never a silent patch. Export lines, Hinglish freeform (Atul ne 220 udhaar bheje aniket ko), and mixed pastes all go through the same path.
  • Python owns money. Totals, equal shares, and minimal who-pays-whom settlements are pure code. The model never does arithmetic. Members come from the model's own payers β€” nobody types a roster.

My hardware is the honest part of this story: a GTX 1650 with 4 GB VRAM. Ollama offloads
30 of 36 layers to the GPU and the rest runs on CPU β€” roughly 10–20 seconds per line.
Not fast, but fully local, and the edit-and-settle table means you verify everything
before money is computed.

Along the way I deleted a whole regex fallback engine. It felt responsible to have a
"backup parser" β€” but a silent worse parser producing wrong money splits is worse than no
parser at all. Now if Ollama is down, you get one clear error with the exact commands to
start it. Fail fast beat fallback.

Why Does Open Innovation Matter?

Four things a closed API physically cannot give this project:

  1. Privacy where it counts. Grocery chats reveal food habits, routines, and who owes whom β€” a sketch of your household finances. That should never sit on someone else's server just to split sabji money. Here it never leaves the laptop.
  2. Zero marginal cost. Month-end splits happen forever; metered tokens would tax every one. This costs rupees zero after a one-time 4.3 GB download.
  3. Offline by construction. Hostel WiFi dies the night before settlement. The split doesn't.
  4. Real control. Apache 2.0 weights I can swap with one .env line (gemma4:e2b-it-qat ↔ gemma4:e2b-it-q4_K_M), prompt, and ground with deterministic code around it. Try changing how a hosted chatbot parses Sabji+jeera+maggi --> 40 + 30.

The open approach didn't just match the closed one β€” for this job it won on every axis
the challenge cares about.

Prize Categories

  • Best Use of Gemma (Gemma 4 E2B quantized, local via Ollama β€” the entire parsing core)

Teammate:

Top comments (0)