DEV Community

Nayananshu Garai
Nayananshu Garai

Posted on

Bill-Bhasa

Hacktoberfest Weekend Challenge: Build for a Friend Submission ๐Ÿค

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

BillBhasha โ€” a privacy-first bill scanner that reads any paper receipt, prescription, or invoice out loud in my parents' own language.

My mother is 62, hindi-first, and has low vision. She still handles every household bill, medicine strip, and ration document by hand. Right now that means either squinting at faded paper or asking someone else to read it for her โ€” and neither option lets her stay independent.

BillBhasha lets her simply take a photo of any document. The app cleans the image, reads the text with open-source OCR, understands it with an open-weight language model, and then speaks the explanation back in warm, simple Hindi or Bengali. It also watches her spending history and speaks a gentle warning when an amount jumps unusually.

Everything runs inside a single free-tier web service. No app store, no login, no credit card. She opens the site, snaps a photo, and listens.

Demo

  • Live link: https://billbhasha.onrender.com (first visit may take ~40s to wake the free service)
  • How it feels for her: open the page โ†’ tap the camera โ†’ snap a bill โ†’ hear a 30-second Hindi explanation โ†’ see a plain โ‚น amount and a simple "check this" warning if something looks off

Code

  • GitHub repo: https://github.com/N-Garai/Bill-Bhasa
  • Stack: FastAPI backend ยท Tesseract OCR ยท SmolLM2-360M (GGUF, Apache-2.0) ยท Piper Hindi/Bengali voice ยท Neon Postgres ยท vanilla JS frontend with GSAP + Three.js

How I Built It

BillBhasha is not a wrapper around a closed API.

Open-weight model at the core

  • The default brain is SmolLM2-360M in GGUF format (LLM_MODEL env var). It turns raw OCR text into a simple Hindi/Bengali explanation with amount, date, and action items.
  • It runs locally inside the same container via llama.cpp subprocess isolation, so there is no external inference bill and no data leaves the box.
  • The model is swappable in one config line: Qwen2.5-0.5B, Phi-3, or any GGUF-compatible model. I started with Phi-3 Mini but it OOM'd on the 0.1 CPU free tier, which is why SmolLM2-360M became the default.

Open harness + local inference

  • FastAPI + Uvicorn serve the app. Tesseract (pytesseract) does the OCR. Both are open source.
  • Heavy work runs in subprocesses with a sequential in-process lock, keeping peak memory under 512 MB on a single CPU.
  • Piper provides an offline Hindi/Bengali voice on-device, so the app still speaks when the network is down.

Open vision fallback

  • When Tesseract returns nothing (faded print, bad light), Gemma (gemma-4-26b-a4b-it) reads the photo directly through the free Google AI Studio API and returns structured JSON.
  • The vision prompt is locked to the user's chosen language, and the response is schema-validated before it reaches the TTS stage.

Why this architecture

  • My parents' data never needs to leave the server. On the SQLite profile, the app runs fully offline.
  • Every behavior swaps via environment variables: LLM_PROVIDER, LLM_MODEL, TESS_LANG, ELEVENLABS_API_KEY, GEMINI_API_KEY. No code changes required.
  • The monthly receipt is $0: Render free web service + Neon free Postgres + zero closed-API spend. ## Why Does Open Innovation Matter? <!-- Why does open innovation matter for what you built? What did it make possible that a closed API wouldn't? --> For my parents, "open" is not a buzzword โ€” it is the only reason this exists at all.

Privacy without trust in a closed server

  • My parents' prescriptions, bank statements, and hospital bills never leave our infrastructure. With SQLite + local inference, the app works with the cable pulled. A closed API would require sending those papers to someone else's server, with no guarantee they won't train on them.

Model swap = accessibility swap

  • Hindi-first, Bengali-first, and English-first users all need the same app, but different voices and different LLMs. Open-weight GGUF models let me swap the brain in one line. Closed APIs would lock me into their pricing, their accents, and their privacy policy.

Cost = access

  • My parents are retired. A $20/month closed-API bill is not a viable long-term plan. SmolLM2-360M, Tesseract, Piper, Render free tier, and Neon free tier together cost $0. Open innovation made this project actually sustainable for the person it was built for.

What failed and what open alternatives made possible

  • Phi-3 Mini: 2.5 GB, OOM on 0.1 CPU. SmolLM2-360M: 360 MB, runs comfortably.
  • EasyOCR: pulled from the box for memory. Tesseract with sparse passes stayed under budget.
  • Render Postgres 30-day reset trap โ†’ Neon free Postgres, no reset.
  • Closed TTS APIs sounded robotic in Hindi/Bengali. ElevenLabs multilingual v2 + Piper on-device gave us a warm, female voice that sounds like home.

Prize Categories

  • Best Use of Render โ€” deployed on Render free web service with Neon Postgres
  • Best Use of Gemma โ€” Gemma 4 vision fallback via Google AI Studio free API (gemma-4-26b-a4b-it) with schema-locked JSON output
  • Best Use of ElevenLabs โ€” ElevenLabs multilingual v2 TTS for Bengali/Hindi with Piper offline fallback

Top comments (0)