This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Promise Keeper is a small, local-first tool designed for a friend with ADHD. It turns voice notes or pasted texts into a short list of concrete commitments, with deadlines and the original words they used.
For someone managing ADHD, keeping track of commitments can take real effort: a promise can be easy to make in conversation and hard to remember later, especially when it's buried in a long, chatty voice note. I wanted to reduce that bit of memory overhead without asking my friend to change how they communicate or maintain yet another complicated system.
Promise Keeper tries to keep the list useful by being conservative. "I'll send you the slides by Friday" is a commitment; "I should probably get to that" and "maybe I'll look at it" are not. Each extracted item keeps its source quote, so my friend can quickly check the context and decide what to do. It's a memory aid, not a treatment or a substitute for tools and strategies that work for the person.
The real app runs locally. It accepts text or audio, transcribes audio on-device, asks a local model to extract commitments, and saves the results to todos.json and todos.md.
Demo
Hosted demo: https://promise-keeper-demo.onrender.com/
The browser demo is a lightweight preview: it processes text in the browser with simple rules, and nothing is uploaded or saved. It does not run the local LLM or transcribe audio; the full CLI app does that on the user's device.
Code
Promise Keeper
Your friend talks fast and says "yeah I'll send that over" or "I'll get you the file by Friday" a dozen times a day - then forgets. This tool listens to their rambly voice notes or texts and pulls out every actual commitment they made, with deadlines, into one clean todo list.
Built for Hacktoberfest "Build for a Friend" (open-source AI at its core).
Why local / open-source matters here
- It's a surveillance-shaped tool. Transcribing someone's voice notes and texts to extract "what you promised" is sensitive by nature - this only works if your friend trusts it never leaves their machine. A cloud API is a non-starter for this use case.
- No internet required. Works on a commute, in a basement office, wherever the rambling happens.
- Free to run constantly. This only has value if it runs on every voice note, all day, forever - an API-metered…
How I Built It
The CLI is three small, swappable stages wired together in main.py: audio or text comes in, a local model turns it into structured commitments, and a plain-file store dedupes and renders them. Everything happens on the machine running it - no API keys, no network calls.
1. Turning a voice note into text
transcribe.py wraps faster-whisper's small model, running on CPU with int8 compute so it stays usable on ordinary hardware:
MODEL_SIZE = "small" # tiny/base = faster, small/medium = better for fast/rambly speech
def _get_model() -> WhisperModel:
global _model
if _model is None:
_model = WhisperModel(MODEL_SIZE, device="cpu", compute_type="int8")
return _model
def transcribe_audio(audio_path: Path) -> str:
model = _get_model()
segments, _info = model.transcribe(str(audio_path), beam_size=5)
return " ".join(segment.text.strip() for segment in segments)
main.py only calls this when the input file extension looks like audio (.m4a, .wav, .mp3, .mp4, .ogg, .flac); plain text files and --text input skip straight to extraction.
2. A deliberately strict extraction prompt
extract_commitments.py sends the transcript to Gemma 3 through Ollama with a system prompt that draws a hard line between a real promise and a vague intention:
SYSTEM_PROMPT = """You extract concrete COMMITMENTS from a transcript of someone talking or texting.
A commitment is something the speaker promised to DO for someone else, optionally with a deadline.
Examples that ARE commitments:
- "I'll send you the deck by Friday" -> task: send the deck, deadline: Friday
...
Examples that are NOT commitments (do not include these):
- "I should really get to that" (vague intention, no real promise)
- "maybe I'll look at it" (hedged, not committed)
...
Return ONLY a JSON array... Each item:
{"task": "...", "deadline": "<as stated, or null>", "source_quote": "<the exact phrase from the transcript>"}
"""
The call runs at temperature: 0.1 to keep output as deterministic as a local model allows, and MODEL = "gemma3" is a single constant, so swapping in whatever Ollama model fits a given machine (or a given friend's speech patterns) is a one-line change.
3. Defensive parsing, because small local models are a little feral
A 3-ish-billion-parameter model run locally doesn't always emit clean JSON. Two small repair steps keep the CLI from crashing on it instead of silently losing commitments:
def _strip_code_fence(text: str) -> str:
text = text.strip()
match = re.match(r"^```
(?:json)?\s*(.*?)\s*
```$", text, re.DOTALL)
return match.group(1) if match else text
def _repair_truncated_array(text: str) -> str:
"""Best-effort fix for local models that emit a stop token before closing
the JSON array: valid objects, missing trailing ']'."""
text = text.rstrip().rstrip(",")
last_close = text.rfind("}")
if last_close != -1:
text = text[: last_close + 1]
if not text.endswith("]"):
text += "]"
return text
If both parses fail, extract_commitments prints a warning with the raw output and returns an empty list rather than raising - a bad model response should never take down the ingest command.
4. Dedup and a human-readable list
todo_store.py treats todos.json as the source of truth and regenerates todos.md on every write. New commitments are deduped against existing ones by (task.lower(), deadline), so re-ingesting the same ramble twice doesn't double the list:
existing_keys = {(t["task"].lower(), t.get("deadline")) for t in todos}
for c in new_commitments:
key = (c["task"].lower(), c.get("deadline"))
if key in existing_keys:
continue
...
todos.md renders as a checklist with each item's source quote underneath it in a blockquote, so the person reviewing it can see the exact words that triggered an entry - not just the model's paraphrase of them.
5. The browser demo doesn't touch an LLM at all
The hosted demo in demo/ is a static page; app.js reimplements a much smaller version of the same idea in plain regex, entirely client-side:
const clauses = text.split(/(?<=[.!?])\s+|\n+|(?:,?\s+and\s+|;\s*)(?=I(?:'ll| will| am going to)\b)/i);
...
if (!promise || /\b(?:maybe|might|should|could|probably|perhaps|if|try to|hope to|not|never)\b/i.test(hedgeCheck)) {
continue; // same conservative instinct as the real prompt, just as a word blacklist
}
It mirrors the real extractor's philosophy - hedged language gets dropped - but it's pattern matching, not a model, so it will miss anything phrased unusually. Nothing typed into it is uploaded or saved; refreshing the page clears it. It exists only so a visitor can feel what the tool does without installing Ollama or Whisper. render.yaml deploys demo/ as a static site on Render, with pushes to main updating it automatically.
The model is intentionally replaceable throughout: swap the Ollama model or tweak the prompt, and the rest of the pipeline doesn't care. That flexibility matters for ADHD-related needs because no two people communicate, organize, or remember in exactly the same way; a fixed workflow shouldn't dictate what counts as useful.
Why Does Open Innovation Matter?
Voice notes and personal messages are sensitive, and that matters even more when someone is using them to support a personal ADHD coping strategy. A tool that analyzes them only works if the person using it trusts where that data goes. Here, the audio transcription and language-model inference happen locally; the transcripts and resulting task list stay on the user's machine instead of being sent to a hosted AI API.
Open-weight models make that privacy boundary practical, while also making the behavior inspectable and adaptable. The extraction prompt is in the repository, the model can be swapped, and a friend can tune what counts as a promise for their own conversations. That flexibility matters because "I got you" might be a real commitment in one person's speech, while "maybe I'll look at it" is not. The goal is not to make one universal ADHD productivity tool, but to let the person using it shape a small aid around their own needs.
Prize Categories
- Best Use of Gemma
- Best Use of GitHub Copilot

Top comments (0)