DEV Community

SRIRAM S
SRIRAM S

Posted on

RecallMate: talk for a minute, get a checklist, with open models running locally

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Most of us have a head full of things we meant to remember: tasks we planned, how the day went, what to do tomorrow. Writing them down takes effort. Saying them out loud takes ten seconds.

RecallMate is a voice journal built on that idea. You talk for a minute and upload the recording. RecallMate transcribes it and turns it into a clean, structured entry: your mood, an energy score from 1 to 10, a checklist of action items with priority and due time, and a short summary. Every entry goes into a timeline archive, and any entry can be exported as Markdown.

It has three modes, so one app can fit different people:

  • Daily log turns a ramble into a to-do list and a snapshot of the day.
  • Study log records what you learned, what's blocking you and what to do next.
  • Health tracks symptoms, sleep, diet and mood.

I built it as a gift for my friend Ragul, so he can talk instead of type. He's asleep as I write this, so he hasn't tried it yet. I'll update this post with his honest reaction as soon as he does.

Voice notes are personal, so RecallMate is designed to keep them out of third-party AI services. Transcription and extraction both run on open models.

Demo

Watch the demo video

In the demo I upload a voice note, click Transcribe & Extract Insights, and the app shows the structured insights, the full transcript and the timeline archive. The live app runs from a Colab notebook through a temporary Cloudflare tunnel whose link changes every run, so I'm sharing a recording.

Code

https://github.com/AWT-SRIRAM/RECALLMATE

The whole project fits in one Colab cell. It installs the dependencies, pulls the model, writes the Streamlit app, runs a smoke test and opens a public link. It also runs locally with Ollama.

How I Built It

Voice note (.wav/.mp3/.m4a)
   -> ffmpeg (decode to 16 kHz mono)
   -> faster-whisper (speech to text, GPU with CPU fallback)
   -> Ollama + Qwen2.5 7B (reads the transcript)
   -> Instructor + Pydantic (forces valid, typed JSON)
   -> SQLite (timeline archive)
   -> Streamlit (upload, insights, transcript, timeline, export)
Enter fullscreen mode Exit fullscreen mode
  • faster-whisper runs OpenAI's open Whisper weights through CTranslate2. The model loads once, tries the GPU with a real warm-up run, and falls back to CPU if the CUDA libraries are missing.
  • Ollama with qwen2.5:7b is the scribe. It runs locally and is called through Ollama's OpenAI-compatible endpoint.
  • Instructor and Pydantic matter a lot with a small local model. Instead of hoping for clean JSON, the output has to validate against a schema (mood, energy, symptoms, action items, summary, tags), and Instructor retries when it doesn't.
  • Journal modes: a sidebar switch swaps the system prompt while the same schema keeps the output structured, so one app serves different needs without separate code paths.
  • Guardrails: the prompt tells the model to use only what's in the transcript, never invent details, and never give medical advice. RecallMate organizes what you said; it doesn't diagnose.
  • Streamlit and SQLite keep the UI and storage simple, with an optional password gate because the app sits behind a public tunnel.
  • One-cell deploy: the runner handles installs, the model pull, a headless smoke test of the audio and schema pipeline, and tunnel startup. It waits until the tunnel link actually responds before showing it, because quick tunnels can print a URL before it's reachable.

Why Does Open Innovation Matter?

A voice journal is about as personal as data gets: how someone slept, how they felt, what's worrying them. With a closed API, every recording and transcript would have to be sent to someone else's servers, and the app would depend on their pricing, limits and uptime.

Open-weight models made a different design possible. Whisper and Qwen run in my own session, so no recording or transcript goes to a third-party AI API, and the same code also runs fully offline on a local machine with Ollama. Entries stay in a local SQLite file, and I can swap the model from the sidebar without touching the code. There's also no per-entry cost, so a friend can journal every day without worrying about usage fees.

Top comments (0)