DEV Community

Aryan Gupta
Aryan Gupta

Posted on

Promise Keeper: Never Forget What You Promised, Even in Hinglish

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Promise Keeper is a small app that remembers what you promised people.

I built it for my friend Amit, who makes a lot of promises in a lot of chats and has forgotten most of them by the end of the day.

The problem isn't making promises. It's that they get buried in messages. A to-do app doesn't help, because nobody stops mid-conversation to write down "told Rahul I'd call tomorrow."

So Promise Keeper works from the message itself:

  • Paste something you said, in English or Hinglish ("Kal subah main tumhe call karunga"). It extracts the promise, who it's owed to, and the deadline.
  • Import a WhatsApp chat export. It reads only your lines and finds promises you probably forgot. "Tomorrow" is resolved relative to when you said it, so a promise from last week correctly shows up as overdue.
  • You confirm everything with date and time pickers before anything is saved.
  • "What am I forgetting?" gives an honest list: overdue first, then coming up, then no deadline. You can also group by person ("what do I owe Rahul?").
  • Draft a nudge for an overdue promise: a short, honest "still on it" message you can edit and send.

Demo

πŸŽ₯ Watch the demo video

I didn't deploy a hosted version on purpose. The app reads private chats, so hosting it would defeat the point. It runs on your own machine in about five minutes (setup is in the README).

Code

πŸ“ Promise Keeper

Built for a friend for the Hacktoberfest Weekend Challenge: Build for a Friend.

A local-first assistant that remembers what you promised people. Paste a message ("I'll send you the report tonight" or "Kal subah main tumhe call karunga"), confirm what it found and later ask "What am I forgetting?"

Promise Keeper

Open-source core: an open-weight model (Gemma 3) running locally through Ollama, with storage in SQLite. No API keys, no accounts, works with Wi-Fi off. Your messages and relationships never leave your machine.

Run it

Requires Python 3.10+ and Ollama.

# 1. Pull a small open-weight model
ollama pull gemma3:4b

# 2. Install and configure
python -m venv .venv && source .venv/bin/activate      # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env                                   # Windows: copy .env.example .env

# 3. Start the app
streamlit run app.py
Enter fullscreen mode Exit fullscreen mode

Settings live in .env (git-ignored). No API keys are…

How I Built It

The AI is Gemma 3 (4B), an open-weight model from Google, running locally through Ollama. The app itself is Streamlit and SQLite. There are no API keys, accounts or remote servers.

A few design decisions matter more than the model choice:

The model never does date maths. Small models are bad at calendars. So the model only copies the phrase ("kal subah", "by Monday"), and plain Python turns it into a real date. That code has 30 test cases, including Hinglish like "5 baje" and "parso shaam".

The "what am I forgetting?" list never comes from the model. It's a database query. A language model asked "what did I promise?" can invent things, which is the worst thing a reminder app can do.

Structured output. The model is constrained to a JSON schema, so I get promises I can parse, not paragraphs.

A human confirms everything. Low-confidence guesses are dropped, and nothing is saved until you check it. Casual chat like "lol nice", "I won't be able to make it" and "if I get time I'll try" should produce nothing, and I wrote test cases for those.

I compared models on 20 labelled messages (13 plain, 3 Hinglish, 4 tricky cases like negation and past tense):

Model Right # of promises Deadlines core edge hinglish Avg s/msg
gemma3:4b 17/20 20/20 11/13 3/4 3/3 8.3
qwen2.5:3b 18/20 20/20 11/13 4/4 3/3 3.6

The two models are effectively tied. With 20 cases, a one-case gap isn't evidence that either is better. Both got every deadline and every Hinglish message right. Qwen was about twice as fast, which matters when importing a long chat.

Both made the same two mistakes: they treated a request ("Can you send me the notes by tonight?") and someone else's action ("Rahul will send the files tomorrow") as my own promises. Gemma also extracted a past-tense statement ("I sent you the report yesterday"). Since both models fail the same way, I read these as prompt weaknesses rather than model weaknesses, and they're the reason every extraction goes through human confirmation. The test set is small and the timings come from one machine, so treat the numbers as rough.

Why Does Open Innovation Matter?

A promise tracker reads your private conversations: who you talk to, what you owe them, what you've been avoiding. That's about as personal as data gets. Amit shouldn't have to send it to a company's servers to get a reminder.

Open weights made three things possible that a closed API wouldn't:

  • Nothing leaves the machine. The data stays on the device by design, not by a privacy policy I'm asking you to trust. The demo video shows it running with Wi-Fi off.
  • It costs nothing to run. There's no per-message bill, so scanning a whole year of a chat export is fine.
  • Swapping models is a one-line change. The same app runs Gemma, Qwen or Llama by changing a name in the sidebar. That let me measure the models against each other on my own cases instead of accepting what one vendor offers.

There's also a language angle. Hinglish, meaning Hindi written in English letters, is how many people actually text. With open models I can test it directly, and eventually fine-tune for it. That's on my list for next steps.

Prize Categories

  • Gemma: Gemma 3 4B is the core extraction model, running locally via Ollama. I benchmarked it against Qwen 2.5 on my own labelled cases, including Hinglish.

Top comments (0)