DEV Community

Vaishnavi Devardekar
Vaishnavi Devardekar

Posted on

RecallBuddy — hackathon submission

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built RecallBuddy for my aunt Marta. She is 71 and has memory trouble.

In a single conversation she will ask the same question three or four times.

I never once found that annoying. But I started noticing how software answers her. Every app answers the fourth time the way you answer someone who is being slow — not cruelly, just with a slight flattening. And the well-meaning ones say it out loud:

"As I mentioned earlier…"
"Like I already told you…"

That is the response that makes her stop asking. And a woman who stops asking stops finding out things.

So my requirement was never "answer well". It was:

Give the fourth answer as fully as the first. Every time.

Not a nicer prompt. An actual guarantee.

RecallBuddy is a small always-on program on a laptop that listens, thinks and speaks. It never says she asked before.

Marta      : When does the bakery open?
RecallBuddy: The bakery on Elm Street opens at 7 in the morning.

Marta      : Sorry, when does the bakery open?
RecallBuddy: The bakery on Elm Street opens at 7 in the morning.

Marta      : What time does the bakery open?
RecallBuddy: The bakery on Elm Street opens at 7 in the morning.

Marta      : Who takes the bins out?
RecallBuddy: Your neighbour Marek comes to take the bins every Wednesday.

Marta      : When do I take my tablets?
RecallBuddy: You take your blood pressure tablets at 8am and 8pm.
Enter fullscreen mode Exit fullscreen mode

0.5–2 seconds per turn on a laptop CPU. It knows about her dog Pepper, her tablets, her bins, her neighbour Elena — because those facts live in a text file she can edit herself.

It runs on her family's laptop. It works with the wifi off. Nothing is sent anywhere, and it costs nothing per conversation.

Demo

This project has no deployed link, and that is the point.

RecallBuddy is offline-first by design. The entire premise is that my aunt's household, her medication schedule and her voice never touch a third party's server. There is no URL in the request path — a test enforces that. So there is nothing to deploy and nothing to sign up for.

What I can show you instead is harder to fake: proof it works with no internet at all.

verify_offline.sh does not tell me to turn the wifi off and trust the output. It builds a Linux network namespace containing nothing but loopback, starts Ollama inside it, and runs the whole pipeline there:

1. Is there an internet?
  PASS  cannot reach https://example.com
  PASS  cannot reach https://api.openai.com
  PASS  cannot reach https://huggingface.co
       iface: lo:

2. Start Ollama inside the namespace
  PASS  Ollama is up on loopback with llama3.2:3b

3. Whisper weights, read from disk (no download)
  PASS  whisper tiny.pt already cached (73M)

4. Full loop: hear -> think -> speak, with no internet
  PASS  the pipeline ran to completion
  PASS  3 repeat questions detected and all answered in full
  PASS  no reply tripped the pattern filter or the audit
  SKIP  speech not exercised here

RESULT: 8 passed, 0 failed
Enter fullscreen mode Exit fullscreen mode

Anything that passes in there cannot have phoned home, because there was no route out to phone home with.

The single SKIP is real and I am reporting it rather than hiding it: WSL2 has no PortAudio and no espeak, so the microphone and speaker were never exercised on this machine. Windows supplies both. Everything else, including the entire AI loop, was verified.

Code

Repository: https://github.com/vaishud12/RecallBuddy-1

microphone ──► whisper tiny ──► llama3.2:3b ──► pyttsx3
 (PortAudio)   (local weights)  (via Ollama)   (SAPI5 / espeak)
                    │               │
              memory_bank.txt    the reply guard
              (a plain .txt)     (plain Python)
Enter fullscreen mode Exit fullscreen mode

Every arrow is a loopback socket or a local file.

File What it does
app.py wiring, CLI, session log, repeat detection, re-ask
companion.py the patient prompt, the guard, the audit
audio_io.py microphone, WAV encode/decode, Whisper
brain.py Ollama client over loopback
speech.py pyttsx3 output
memory.py knowledge-bank parser
memory_bank.txt what it knows about Marta — editable by her
verify_offline.sh the network-namespace proof

Run it:

ollama pull llama3.2:3b
python -m venv .venv
.venv/bin/pip install torch --index-url https://download.pytorch.org/whl/cpu
.venv/bin/pip install -r requirements.txt
.venv/bin/python app.py --name Marta
Enter fullscreen mode Exit fullscreen mode

Windows: double-click run_windows.bat. Linux/macOS: ./run.sh.

How I Built It

Open-source AI used: openai-whisper (tiny) for speech-to-text, llama3.2:3b served locally through Ollama, and pyttsx3 for speech output. Everything runs on CPU.

The interesting engineering is not the plumbing. It is that patience has to be a guarantee, not a request.

With a hosted API, being kind is something you ask for. You write "always be patient, never imply the user repeated themselves" into the system prompt and take what comes back. You are hoping. And the hope is invisible — you cannot grep a response for a property you did not get.

"Usually patient" is not good enough here. The failure mode is not a bad sentence. It is a person deciding she is a burden.

Running the inference myself means the property lives in code I own. Twenty-three regexes sit between the model and the speaker:

# companion.py
_REMINDER_REST = r"[^.!?;,]*"     # a clause, never a whole sentence

_YOU = r"\byou(?:"
        r"\s*'?(?:ve|d|re|ll|m)\b"           # you've, you'd, you're, you'll
        r"|\s+(?:will\s+have|have|had|are|were|will)\b"
        r"|\s+"                               # bare "you asked"
        r")\s*"

FORBIDDEN = [
    rf"\bas (?:i |we )?(?:already |just )?(?:said|told|mentioned|asked)\b{_REMINDER_REST}",
    rf"{_YOU}(?:mentioned|said|told me|noted|brought up)\b{_REMINDER_REST}",
    rf"{_YOU}(?:just |already )?(?:asked|said|told me)\b{_REMINDER_REST}",
    # ... 23 patterns total
]
Enter fullscreen mode Exit fullscreen mode

"you already asked" cannot reach her, because the code refuses to emit it.

That one character class — [^.!?;,]* — is the whole difference between removing a reminder and removing the answer:

"You mentioned earlier it was 7am, so she goes there in the morning."
                            ↓  stop at the comma
"So she goes there in the morning."
Enter fullscreen mode Exit fullscreen mode

Without the comma in that class the pattern eats through to the full stop, the reply becomes ".", and pyttsx3 reads that aloud as a pause. I got this wrong first and caught it in a live run.

Four layers, because one was not enough

1 · Clause-scoped patterns — as above. Delete the reminder, keep the answer.

2 · Two passes. The filter runs before and after the tidy-up. This is not defensiveness, it is a real bug: "I already said that." gets tidied by stripping the trailing that as a dangling word, producing "I already said." — which is itself a forbidden phrase. The guard was manufacturing the exact thing it exists to prevent.

3 · An independent, blunt audit. A vocabulary check, no regex, deliberately not reusing the patterns above. This matters: my live test asserted "no forbidden phrase in the reply" by running FORBIDDEN_RE over the reply — the filter was being tested with itself. It passed while a real leak sat in the transcript:

You've mentioned that before, but I'll always be happy to remind you.

Every pattern required a literal space after you. "You've" has an apostrophe, so all 23 missed it — and so did the test, because the test used the same 23.

4 · Ask again. If the guard ends up with nothing usable it sets a bailed flag and the app re-asks once. Being kind is not enough if we have also stopped answering.

The bug I would most like feedback on

The audit is a blunt instrument, and blunt instruments fire on the wrong thing.

My first version treated once and twice as reminder vocabulary. In a live session it deleted:

"You take them once in the morning and once in the evening."

…and replaced a medication schedule with a contentless line. For a woman whose blood pressure tablets are the reason the app exists, that is the worst failure available — worse than saying "as I mentioned", because now she has been told nothing at all.

The fix was to split the audit by whether a term has an innocent use in a factual answer:

AUDIT_TERMS_STRICT = {"already", "mentioned", "repeated", "looping"}
AUDIT_TERMS_SOFT   = {"before", "again", "said", "keep", "remind", ...}
Enter fullscreen mode Exit fullscreen mode

Strict terms can trigger a bail. Soft terms only warn, because "the shop closes at six, so go before four" is perfectly good English.

Is "does this word have an innocent use in a factual answer" the right line to draw? It works, but it feels hand-tuned. An entailment check, a cheap classifier, or a smarter fallback that does not have to choose would all be better.

The knowledge bank is a .txt

## Routines and facts worth holding on to
- Her dog is named Pepper. Pepper is 12 and sleeps on the blue blanket.
- She takes her tablets at 8am and 8pm. The blue ones are for blood pressure.

## Things to be careful about
- Do not bring up the car. She is still sad about it.
Enter fullscreen mode Exit fullscreen mode

No vector database, deliberately: she can edit it (Notepad, no tooling), she can read it (she can see exactly what the machine thinks it knows about her), and it fits (25 facts into a 3B prompt without swamping it).

Tests

58 tests — 52 fast, 6 against the real local model. 30 hostile phrasings, each asserted to be stripped and to keep its answer and never to come back empty. Plus a parallel set of innocent replies asserted to pass through completely untouched, because a guard that eats answers is just a quieter failure.

Why Does Open Innovation Matter?

Because for this project, "usually patient" is a failure and only a guarantee is good enough.

Running the weights myself is what makes the guarantee possible. The rest of the open-based wins are concrete:

closed API this
Offline needs a network proven in a namespace with no route out, 8/8
Her data sent to a vendor, kept per their policy never leaves the machine
A behaviour I dislike a habit I cannot patch a regex I can read
Swapping models whatever they permit llama3.2:3b → llama3.1:8b → qwen2.5, guard unchanged
Cost per token, forever one download, then zero
If the vendor shuts down it is over it still runs

Two of those I want to underline.

Privacy is not decorative here. This is software for a woman whose cognitive decline is the most private thing about her, and whose family is already anxious about her. Shipping her household and medication schedule to a third party in order to fix a politeness problem is a bad trade. Open weights meant I did not have to make that trade.

A text file is a better interface than an API for the person it is about. She can open memory_bank.txt in Notepad and add a line about her dog herself. No account, no key, no index to rebuild. Closed systems do not let the person being cared for have that.

My Agent Session

I could not include a DevRelay session — DevRelay was not available in my environment, and I would rather say so than fake it.

What I have instead is the verification trail: 58 tests, 6 of which drive the real local model, a network-namespace offline proof, and the failure transcripts quoted above. The bugs I wrote about in this post were found by running things rather than by reasoning about them, and the raw outputs are in the repo.

What It Cannot Do

I would rather be blunt about this than have someone find out later.

  • It is a 3B model and it shows. It invents details. It once said Marek collects the bins "at the bakery". It once answered "every day" when the bank says Tuesdays and Thursdays. It has contradicted itself inside a single session. The knowledge bank is small and true; the model's additions to it are not reliable enough for anything that matters.
  • Third person slips through despite the prompt — it will say "she goes there" instead of "you go there". I chose not to rewrite pronouns in code, because "she" is often correct (Elena, the neighbour) and blind substitution would be worse than the slip.
  • Whisper tiny is weak on a quiet room with an older voice. base is noticeably better and still fast on CPU.
  • Not a medical device. The prompt refuses medical, legal and financial questions.
  • Marta has never actually used it. It is verified against her questions in simulation, on my machine. She has not had it in her hands yet. That is the limitation I am working on next, and it is the reason I am asking for criticism rather than applause.

Prize Categories


Questions for me

I would genuinely like to be challenged on the guard design, and on that strict/soft audit split in particular. If you think there is a more principled way to decide when a sentence is too damaged to speak, I want to hear it.

Leave a comment and I will answer.


MIT licensed.

Top comments (0)