<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Debashis Nayak</title>
    <description>The latest articles on DEV Community by Debashis Nayak (@deb2000sudo).</description>
    <link>https://dev.to/deb2000sudo</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1403847%2Faf7b5380-6545-4d9b-9921-2bb36e4889fc.png</url>
      <title>DEV Community: Debashis Nayak</title>
      <link>https://dev.to/deb2000sudo</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/deb2000sudo"/>
    <language>en</language>
    <item>
      <title>MemoTask: a voice to-do for a friend who struggles to type</title>
      <dc:creator>Debashis Nayak</dc:creator>
      <pubDate>Mon, 05 Oct 2026 06:47:47 +0000</pubDate>
      <link>https://dev.to/deb2000sudo/memotask-a-voice-to-do-for-a-friend-who-struggles-to-type-pjk</link>
      <guid>https://dev.to/deb2000sudo/memotask-a-voice-to-do-for-a-friend-who-struggles-to-type-pjk</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;A friend of mine has a birth hand difference. A keyboard is a slow, painful way to capture a thought, so small tasks never make it onto a list.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MemoTask&lt;/strong&gt; is a voice-first to-do for that person. They hold a microphone, speak the way they already think, and get back a structured task: a title, details, a priority, and a due time. Nothing is saved until they look at it and confirm.&lt;/p&gt;

&lt;p&gt;The open piece is the reasoning. Task understanding runs on &lt;strong&gt;Gemma 3 4B&lt;/strong&gt; through Ollama on the same computer. Speech-to-text is &lt;strong&gt;ElevenLabs Scribe v2&lt;/strong&gt;, and the app says so on the screen. I did not want a local model pretending it had heard the audio.&lt;/p&gt;

&lt;p&gt;There is also &lt;strong&gt;Remix&lt;/strong&gt;. On any saved task they can ask Gemma to break it into steps. The suggestion sits next to the original. The todo changes only after &lt;strong&gt;Apply Changes&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;There is no hosted demo. The reasoning model is on the laptop, and a public server would hide the thing the project is about. Here is the path I actually ran:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;ollama pull gemma3:4b&lt;/code&gt;, then &lt;code&gt;npm run dev&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Hold the microphone and say something ordinary, like "buy milk tomorrow morning."&lt;/li&gt;
&lt;li&gt;The screen shows ElevenLabs transcribing, then "Gemma is reading the transcript on this computer."&lt;/li&gt;
&lt;li&gt;A draft appears. Edit it, confirm it, or throw it away. Confirm is the only way it lands in IndexedDB.&lt;/li&gt;
&lt;li&gt;On a saved task, click &lt;strong&gt;Remix&lt;/strong&gt;, optionally add a note, and wait for the local suggestion. Cancel leaves the todo alone. Apply writes the new title, description, priority, category, and subtasks. The due time, the recording, and the transcript stay put.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The header stays honest the whole time: "Voice transcription uses ElevenLabs. Task understanding runs locally using Gemma 3 4B." A status pill reads &lt;strong&gt;Gemma 3 4B • Local&lt;/strong&gt; or &lt;strong&gt;Gemma 3 4B • Offline&lt;/strong&gt;. Offline does not secretly switch the thinking to a closed model. Transcription still needs ElevenLabs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/deb2000-sudo/memo-Task" rel="noopener noreferrer"&gt;https://github.com/deb2000-sudo/memo-Task&lt;/a&gt;&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/deb2000-sudo" rel="noopener noreferrer"&gt;
        deb2000-sudo
      &lt;/a&gt; / &lt;a href="https://github.com/deb2000-sudo/memo-Task" rel="noopener noreferrer"&gt;
        memo-Task
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;MemoTask&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;MemoTask is a voice-first to-do. You speak a thought, ElevenLabs Scribe v2 transcribes it, and Gemma 3 4B (running locally through Ollama) turns the transcript into a task. Nothing is saved until you confirm it. Saved tasks stay in this browser’s IndexedDB.&lt;/p&gt;
&lt;p&gt;A saved task can also be remixed: Gemma proposes a clearer title, description, priority, category, and subtasks. The original task changes only after you click &lt;strong&gt;Apply Changes&lt;/strong&gt;.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What you need&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://nodejs.org" rel="nofollow noopener noreferrer"&gt;Node.js&lt;/a&gt; 20.9 or newer, with npm&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ollama.com" rel="nofollow noopener noreferrer"&gt;Ollama&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;An &lt;a href="https://elevenlabs.io" rel="nofollow noopener noreferrer"&gt;ElevenLabs&lt;/a&gt; API key (speech-to-text only)&lt;/li&gt;
&lt;li&gt;A microphone, and a browser that can record audio (Chrome, Edge, or Safari)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Task understanding does not use a hosted chat model. Gemma runs on your machine at &lt;code&gt;http://127.0.0.1:11434&lt;/code&gt;.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;1. Get the code&lt;/h2&gt;
&lt;/div&gt;
&lt;div class="highlight highlight-source-shell notranslate position-relative overflow-auto js-code-highlight"&gt;
&lt;pre&gt;git clone https://github.com/deb2000-sudo/memo-Task.git
&lt;span class="pl-c1"&gt;cd&lt;/span&gt; memo-Task&lt;/pre&gt;

&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;2. Install the app&lt;/h2&gt;

&lt;/div&gt;
&lt;div class="highlight highlight-source-shell notranslate position-relative overflow-auto js-code-highlight"&gt;
&lt;pre&gt;npm install&lt;/pre&gt;

&lt;/div&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;3. Install and start Ollama&lt;/h2&gt;

&lt;/div&gt;
&lt;p&gt;Install Ollama from &lt;a href="https://ollama.com" rel="nofollow noopener noreferrer"&gt;ollama.com&lt;/a&gt;, then start it.&lt;/p&gt;
&lt;p&gt;On macOS, opening the Ollama…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/deb2000-sudo/memo-Task" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The browser never talks to Ollama, and it never sees the ElevenLabs key.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
  mic[Microphone] --&amp;gt; next[Next.js server]
  next --&amp;gt; scribe[ElevenLabs Scribe v2]
  scribe --&amp;gt; transcript[Transcript]
  transcript --&amp;gt; gemma[Gemma 3 4B via Ollama]
  gemma --&amp;gt; preview[Preview]
  preview --&amp;gt; person[User confirms]
  person --&amp;gt; db[IndexedDB]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Remix uses the same model, through a different prompt, and it does not run until the button is clicked:&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
  todo[Saved todo] --&amp;gt; dialog[Remix dialog]
  dialog --&amp;gt; api["POST /api/ai/remix-task"]
  api --&amp;gt; gemma[Gemma 3 4B]
  gemma --&amp;gt; schema[Schema check]
  schema --&amp;gt; compare[Current vs suggestion]
  compare --&amp;gt; apply[Apply Changes]
  apply --&amp;gt; db[IndexedDB]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Stack: Next.js, TypeScript, Tailwind, shadcn/ui, Dexie for IndexedDB, the ElevenLabs SDK for transcription only, and Ollama at &lt;code&gt;127.0.0.1:11434&lt;/code&gt; with the model name fixed to &lt;code&gt;gemma3:4b&lt;/code&gt;. There is no switch that points task understanding at OpenAI, Gemini, Claude, or any other hosted model.&lt;/p&gt;

&lt;p&gt;Two community posts shaped the design. &lt;a href="https://dev.to/jangwook_kim_e31e7291ad98/ollama-structured-outputs-in-practice-getting-type-safe-json-from-local-llms-with-pydantic-m38"&gt;Jangwook Kim's write-up of Ollama structured outputs&lt;/a&gt; is why the chat request sends a JSON schema in &lt;code&gt;format&lt;/code&gt;, then validates the text anyway and retries once when the JSON is broken. On a 4B model, constrained decoding is not a substitute for a parser. &lt;a href="https://dev.to/zackriya/local-meeting-notes-with-whisper-transcription-ollama-summaries-gemma3n-llama-mistral--2i3n"&gt;Meetily&lt;/a&gt; was the closest product I found: audio in, a local Gemma summary out. MemoTask is the smaller version of that idea for one person's errands, with a confirmation step Meetily's meeting flow does not need. &lt;a href="https://dev.to/sanskarin/building-local-first-ai-apps-what-changes-when-the-data-stays-on-the-device-15mi"&gt;Sanskar's local-first piece&lt;/a&gt;, and the comments under it, is why the status pill names what leaves the device. A badge that only says "local" would be a lie here, because the recording does leave.&lt;/p&gt;

&lt;p&gt;The remix prompt is short on purpose. Gemma is told to keep the original intent, skip invented facts, skip deadlines, and return JSON only:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;REMIX_TASK_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;You are a local task planning assistant.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Your job is to transform an existing task into a clearer, more actionable plan while preserving the user's original intent.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Do not invent facts.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Do not add unnecessary work.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Only create useful actionable subtasks.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Do not create deadlines.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Do not claim knowledge that was not provided.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Return valid JSON only.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Voice drafts and remix plans share one validator. A missing description, a priority outside &lt;code&gt;low | medium | high | null&lt;/code&gt;, or an extra key is rejected. The route then returns the same error shape as the rest of the API: Ollama down, model missing, timeout, empty body, bad JSON.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Open Models Matter Here
&lt;/h2&gt;

&lt;p&gt;The task list is personal. Physio at 4, milk, an interview to prepare for. I did not want that text sent to a hosted chat model just to get a title and a priority.&lt;/p&gt;

&lt;p&gt;Gemma on the laptop is what makes the project work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The thinking stays on the machine.&lt;/strong&gt; The transcript is read by Ollama at &lt;code&gt;127.0.0.1&lt;/code&gt;. Confirmed tasks, transcripts, and recordings live in this browser's IndexedDB. Cancel throws the draft away.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It still runs when the model host is the laptop.&lt;/strong&gt; If Ollama is off, the pill says Offline. The app does not fall through to a closed API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The model can be swapped later&lt;/strong&gt; without touching the screen. The UI calls a task-understanding interface. Today that interface is Gemma 3 4B. A different open model would be a provider change, not a redesign.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reasoning costs nothing per task.&lt;/strong&gt; A 4B model on a laptop is enough to pull a title, a priority, and a few subtasks out of a sentence. A closed model would be a worse fit: more data leaving, and a bill for something a small local model already does.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The closed piece is transcription, and only transcription. ElevenLabs Scribe v2 hears the memo. Gemma never receives the audio file. That split is the honest version of "open at the core": the decision about what the task &lt;em&gt;is&lt;/em&gt; belongs to the open model. A fully local speech model would be better for this friend, and it is the obvious next step. Shipping a fake "fully local" badge would have been worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma.&lt;/strong&gt; Gemma 3 4B through Ollama is the only model that reads a transcript or remixes a task. It proposes the structure. The person still approves it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of ElevenLabs.&lt;/strong&gt; Scribe v2 turns the recording into text so that local Gemma has something to read. The UI names ElevenLabs wherever audio is uploaded.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What They Said
&lt;/h2&gt;

&lt;p&gt;I have not put this in their hands yet, and I am not going to invent a reaction. The app is ready to run on a laptop next to them: one button, a preview, and a list that does not change unless they say so.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;A local speech model, so the recording can stay on the device too. Until then the privacy panel will keep saying that the audio leaves and the understanding does not.&lt;/p&gt;

&lt;p&gt;This article was drafted with AI assistance, then checked against the app I actually ran.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
