DEV Community

Debashis Nayak
Debashis Nayak

Posted on AI-assisted

MemoTask: a voice to-do for a friend who struggles to type

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

A friend of mine has a birth hand difference. A keyboard is a slow, painful way to capture a thought, so small tasks never make it onto a list.

MemoTask is a voice-first to-do for that person. They hold a microphone, speak the way they already think, and get back a structured task: a title, details, a priority, and a due time. Nothing is saved until they look at it and confirm.

The open piece is the reasoning. Task understanding runs on Gemma 3 4B through Ollama on the same computer. Speech-to-text is ElevenLabs Scribe v2, and the app says so on the screen. I did not want a local model pretending it had heard the audio.

There is also Remix. On any saved task they can ask Gemma to break it into steps. The suggestion sits next to the original. The todo changes only after Apply Changes.

Demo

There is no hosted demo. The reasoning model is on the laptop, and a public server would hide the thing the project is about. Here is the path I actually ran:

  1. ollama pull gemma3:4b, then npm run dev.
  2. Hold the microphone and say something ordinary, like "buy milk tomorrow morning."
  3. The screen shows ElevenLabs transcribing, then "Gemma is reading the transcript on this computer."
  4. A draft appears. Edit it, confirm it, or throw it away. Confirm is the only way it lands in IndexedDB.
  5. On a saved task, click Remix, optionally add a note, and wait for the local suggestion. Cancel leaves the todo alone. Apply writes the new title, description, priority, category, and subtasks. The due time, the recording, and the transcript stay put.

The header stays honest the whole time: "Voice transcription uses ElevenLabs. Task understanding runs locally using Gemma 3 4B." A status pill reads Gemma 3 4B • Local or Gemma 3 4B • Offline. Offline does not secretly switch the thinking to a closed model. Transcription still needs ElevenLabs.

Code

https://github.com/deb2000-sudo/memo-Task

MemoTask

MemoTask is a voice-first to-do. You speak a thought, ElevenLabs Scribe v2 transcribes it, and Gemma 3 4B (running locally through Ollama) turns the transcript into a task. Nothing is saved until you confirm it. Saved tasks stay in this browser’s IndexedDB.

A saved task can also be remixed: Gemma proposes a clearer title, description, priority, category, and subtasks. The original task changes only after you click Apply Changes.

What you need

  • Node.js 20.9 or newer, with npm
  • Ollama
  • An ElevenLabs API key (speech-to-text only)
  • A microphone, and a browser that can record audio (Chrome, Edge, or Safari)

Task understanding does not use a hosted chat model. Gemma runs on your machine at http://127.0.0.1:11434.

1. Get the code

git clone https://github.com/deb2000-sudo/memo-Task.git
cd memo-Task
Enter fullscreen mode Exit fullscreen mode

2. Install the app

npm install
Enter fullscreen mode Exit fullscreen mode

3. Install and start Ollama

Install Ollama from ollama.com, then start it.

On macOS, opening the Ollama…

How I Built It

The browser never talks to Ollama, and it never sees the ElevenLabs key.

flowchart LR
  mic[Microphone] --> next[Next.js server]
  next --> scribe[ElevenLabs Scribe v2]
  scribe --> transcript[Transcript]
  transcript --> gemma[Gemma 3 4B via Ollama]
  gemma --> preview[Preview]
  preview --> person[User confirms]
  person --> db[IndexedDB]

Remix uses the same model, through a different prompt, and it does not run until the button is clicked:

flowchart LR
  todo[Saved todo] --> dialog[Remix dialog]
  dialog --> api["POST /api/ai/remix-task"]
  api --> gemma[Gemma 3 4B]
  gemma --> schema[Schema check]
  schema --> compare[Current vs suggestion]
  compare --> apply[Apply Changes]
  apply --> db[IndexedDB]

Stack: Next.js, TypeScript, Tailwind, shadcn/ui, Dexie for IndexedDB, the ElevenLabs SDK for transcription only, and Ollama at 127.0.0.1:11434 with the model name fixed to gemma3:4b. There is no switch that points task understanding at OpenAI, Gemini, Claude, or any other hosted model.

Two community posts shaped the design. Jangwook Kim's write-up of Ollama structured outputs is why the chat request sends a JSON schema in format, then validates the text anyway and retries once when the JSON is broken. On a 4B model, constrained decoding is not a substitute for a parser. Meetily was the closest product I found: audio in, a local Gemma summary out. MemoTask is the smaller version of that idea for one person's errands, with a confirmation step Meetily's meeting flow does not need. Sanskar's local-first piece, and the comments under it, is why the status pill names what leaves the device. A badge that only says "local" would be a lie here, because the recording does leave.

The remix prompt is short on purpose. Gemma is told to keep the original intent, skip invented facts, skip deadlines, and return JSON only:

export const REMIX_TASK_PROMPT = [
  "You are a local task planning assistant.",
  "Your job is to transform an existing task into a clearer, more actionable plan while preserving the user's original intent.",
  "Do not invent facts.",
  "Do not add unnecessary work.",
  "Only create useful actionable subtasks.",
  "Do not create deadlines.",
  "Do not claim knowledge that was not provided.",
  "Return valid JSON only.",
].join(" ");
Enter fullscreen mode Exit fullscreen mode

Voice drafts and remix plans share one validator. A missing description, a priority outside low | medium | high | null, or an extra key is rejected. The route then returns the same error shape as the rest of the API: Ollama down, model missing, timeout, empty body, bad JSON.

Why Open Models Matter Here

The task list is personal. Physio at 4, milk, an interview to prepare for. I did not want that text sent to a hosted chat model just to get a title and a priority.

Gemma on the laptop is what makes the project work:

  • The thinking stays on the machine. The transcript is read by Ollama at 127.0.0.1. Confirmed tasks, transcripts, and recordings live in this browser's IndexedDB. Cancel throws the draft away.
  • It still runs when the model host is the laptop. If Ollama is off, the pill says Offline. The app does not fall through to a closed API.
  • The model can be swapped later without touching the screen. The UI calls a task-understanding interface. Today that interface is Gemma 3 4B. A different open model would be a provider change, not a redesign.
  • Reasoning costs nothing per task. A 4B model on a laptop is enough to pull a title, a priority, and a few subtasks out of a sentence. A closed model would be a worse fit: more data leaving, and a bill for something a small local model already does.

The closed piece is transcription, and only transcription. ElevenLabs Scribe v2 hears the memo. Gemma never receives the audio file. That split is the honest version of "open at the core": the decision about what the task is belongs to the open model. A fully local speech model would be better for this friend, and it is the obvious next step. Shipping a fake "fully local" badge would have been worse.

Prize Categories

  • Best Use of Gemma. Gemma 3 4B through Ollama is the only model that reads a transcript or remixes a task. It proposes the structure. The person still approves it.
  • Best Use of ElevenLabs. Scribe v2 turns the recording into text so that local Gemma has something to read. The UI names ElevenLabs wherever audio is uploaded.

What They Said

I have not put this in their hands yet, and I am not going to invent a reaction. The app is ready to run on a laptop next to them: one button, a preview, and a list that does not change unless they say so.

What's Next

A local speech model, so the recording can stay on the device too. Until then the privacy panel will keep saying that the audio leaves and the understanding does not.

This article was drafted with AI assistance, then checked against the app I actually ran.

Top comments (0)