DEV Community

Taqui
Taqui

Posted on AI-assisted

Buy This, Not That, Drop This There” - I Made It Make Sense

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

The short version

Someone at home gives you a rapid-fire list while you are already halfway out the door:

“Bring 1 litre milk. Don’t buy bread—we already have it. Drop the parcel at the post office. And get three notebooks.”

Most task apps make the person receiving that message do all the work: stop, type, split, remember, and interpret the cancellation correctly.

Nikalll turns that messy family request into a reviewed pocket card. You can speak or paste the request, let an open-weight Gemma model propose the errands, review every item, confirm destinations, and carry the final card outside.

The important part is not that AI produced a list. The important part is that cancelled items remain visible, quantities remain attached, uncertainty stays visible, and the person—not the model—approves what leaves the house.

What I Built

Nikal is a voice-first family-errand handoff for the small, real-world jobs that usually arrive as noisy voice notes or mixed Hindi/Hinglish/English messages.

It is for someone who is leaving home and needs to remember what another person asked them to buy, collect, or drop off.

The flow is deliberately short:

  1. Capture — record a short voice note or paste the original message.
  2. Transcribe — ElevenLabs Scribe v2 turns voice into editable text.
  3. Extract — Gemma, called through Backboard, proposes structured errands, quantities, cancellations, destination hints, and clarification questions.
  4. Review & assign — the original request stays visible while the person edits, cancels, defers, or assigns each item.
  5. Pocket card — Nikal groups approved tasks by destination.
  6. Outside — the person marks each item done, unavailable, deferred, or unresolved.

Nikal is not an autonomous shopping agent. It does not buy anything, message anyone, discover shops, track GPS, or pretend that a model's valid JSON means the request was understood.

The failure this is designed to prevent

The most dangerous failure is not a crash. It is a plausible-looking list that silently drops the important part:

  • “Do not buy bread” becomes “buy bread.”
  • “Two litres” loses its quantity.
  • “Drop this at the post office” loses its destination.
  • An unclear shop becomes a confident guess.
  • An unavailable item is quietly counted as completed.

Nikal treats these as product states, not edge cases.

What the card preserves

  • Included errands.
  • Explicitly cancelled errands.
  • Deferred or unresolved errands.
  • Quantities and units.
  • The user's original source text.
  • Destination assignments and manual corrections.
  • Honest completion outcomes.

The prepared workspace is saved in versioned browser storage. The AI request is online and hosted; the prepared local card does not require another model call to mark outcomes. I do not claim that this prototype is a fully offline AI application or a verified installable PWA.

Demo

Live demo: nikalll.onrender.com

Video demo:

The video walkthrough shows the complete story rather than a feature tour:

  1. Start with a messy family request.
  2. Keep the cancellation and quantity visible.
  3. Let Gemma produce a draft.
  4. Correct or confirm the destination in the review step.
  5. Prepare the pocket card.
  6. Mark one task complete and leave an unavailable/deferred state visible.

If the hosted providers are unavailable, the app keeps the source and offers the manual path instead of manufacturing a successful AI result.

Code

🚶 Nikalll (WalkRecall)

Turn chaotic doorstep voice notes into a structured, offline-ready errand card.

Live Demo Next.js 16 Google Gemma ElevenLabs TypeScript Strict

Speak once before leaving home, review in seconds, and walk out with a glanceable neighborhood checklist that works completely offline.

Explore Live Demo · Architecture · Features · Quick Start



💡 The Problem: The "Doorstep Voice Note" Chaos

Every time you step out of the house, your mother, roommate, or partner shouts a rapid-fire, chaotic list of errands:

"Arre suno! Bahar ja rahe ho toh sabzi mandi se 1 kg aloo aur dhaniya le lena... par tamatar mat lana ghar pe already hain... phir kirana store se 2 litre doodh aur dahi le lena, par bread bilkul mat lana... aur chemist se BP syrup le lena, band-aid mat lana wo drawer me hai!"

Why traditional tools fail:

  1. To-Do Apps are too slow: No one has the time or patience to manually type 8 separate items with…




The project is a Next.js + TypeScript application with the domain contracts, provider boundaries, tests, and setup instructions committed alongside the UI.

The most relevant areas are:

  • app/page.tsx — the capture, review, and pocket-card flow.
  • app/api/extract/route.ts — bounded, consented text-extraction boundary.
  • app/api/transcribe/route.ts — bounded audio-upload boundary.
  • lib/server/backboard.ts — explicit Backboard provider/model request.
  • lib/server/elevenlabs.ts — ElevenLabs Speech-to-Text adapter.
  • lib/domain/ — schemas and deterministic card/evidence logic.
  • tests/ — domain and provider-normalization regression tests.

How I Built It

Open-weight AI is the core, not decoration

Gemma performs the central interpretation task: turning conversational, mixed-language household instructions into a bounded extraction draft.

The browser never chooses the provider or model. The server selects them from environment configuration and sends a fixed prompt with:

  • explicit provider and model names;
  • non-streaming output;
  • JSON output requested as a formatting aid;
  • memory disabled;
  • web search disabled;
  • no tools or autonomous actions.

The default repository configuration uses Gemma through the openrouter provider in Backboard. The exact model is server-configurable; the repository default is google/gemma-3-27b-it. No API key is sent to the client.

Gemma's output is not trusted just because it is JSON. Zod validates the response, and the server checks that every evidence quote occurs verbatim in the original source. The model cannot return trusted destination IDs, application IDs, timestamps, completion states, or approval decisions.

ElevenLabs is a focused input boundary

Voice is useful because the original problem often starts as speech. ElevenLabs Scribe v2 handles transcription; Nikal then shows the transcript as editable text before extraction.

Audio uploads are size-bounded, accepted only through the server route, and not stored as application data. The UI can always fall back to text entry.

Deterministic code owns trust-sensitive decisions

Ordinary TypeScript—not the model—owns:

  • schema validation;
  • evidence matching;
  • destination alias matching;
  • inclusion/cancellation/deferment state;
  • card preparation rules;
  • local persistence;
  • outcome tracking.

An included item without a confirmed destination cannot enter the prepared card. An unresolved item must be resolved or explicitly deferred. An unavailable item is never silently marked done.

Stack

  • Next.js 16 App Router
  • React 19 and strict TypeScript
  • Tailwind CSS 4 with the existing shadcn/Base UI setup
  • Hugeicons
  • Zod for runtime contracts
  • Gemma via Backboard for extraction
  • ElevenLabs Scribe v2 for transcription
  • Render for the public deployment

What I Learned From the First Attempt

My previous hackathon project taught me that a polished interface and a long feature list are not evidence that the core problem was solved.

So this project is intentionally narrower. The demo has one complete story: a messy request, a source-linked AI draft, a human review, a pocket card, and an honest outcome. I kept the original text visible because an extraction that looks correct can still be wrong. I kept cancellation and uncertainty as first-class states because hiding them is worse than showing an incomplete result.

I also removed claims that I could not prove:

  • hosted Gemma is not described as offline AI;
  • browser persistence is not described as a fully offline PWA;
  • a model's JSON is not described as approval;
  • destination labels are not described as live shop discovery;
  • no accuracy percentage is claimed without a labelled, held-out evaluation;
  • an unavailable errand is not converted into a success state.

The most useful engineering decision was to make provider output cross a small, boring boundary: validate it, verify its evidence against the original message, then let deterministic application code decide what the user can approve.

Why Does Open Innovation Matter?

The open model is useful here because Nikal is not trying to replace a person with an opaque assistant. It is trying to make a messy human handoff inspectable and correctable.

With an open-weight model at the center of the extraction step:

  • the model and provider can be changed without rewriting the product workflow;
  • the extraction prompt and JSON contract can be inspected;
  • failures can be compared against a manual baseline instead of hidden behind a chat interface;
  • the application can keep approval, evidence, and completion outside the model;
  • a future self-hosted Gemma deployment could reduce the hosted-data boundary for sensitive household messages;
  • the same safety contract can validate different open models.

Backboard gives the application an explicit model boundary rather than a browser-side vendor dependency. I can select the Gemma provider/model on the server, turn memory and web search off, and keep the rest of the app independent of that choice.

ElevenLabs is deliberately smaller in scope: it solves the voice-capture problem, then hands editable text to the same review pipeline. It does not become a second decision-maker.

There is an important caveat: the current public deployment uses hosted services. Open innovation makes the model layer replaceable and inspectable; it does not magically make hosted inference private. Nikal tells the user when a request is sent for transcription or extraction.

My Agent Session

I built Nikal with OmniRush.ai as a coding and review partner. The agent helped with the architecture, domain contracts, provider boundaries, failure handling, and test cases; I made the product decisions and kept the final scope focused on one physical-world handoff.

I am not embedding a DevRelay session link in this draft because I do not have a public saved-session URL to link to. This section can be removed before publishing, or replaced with the DevRelay URL if one is created.

Prize Categories

  • Best Use of Gemma — Gemma is responsible for the core interpretation problem: quantities, negations, cancellations, destination hints, and clarification candidates in noisy household language. It is not decorative chat.
  • Best Use of Backboard — Backboard is the explicit server-side model boundary. The request fixes the provider/model, disables memory and web search, requests structured output, and validates the result before it reaches the review UI.
  • Best Use of ElevenLabs — ElevenLabs Scribe v2 handles the short voice-note transcription path and returns editable text before extraction.
  • Best Use of Render — the working Next.js application and its provider-backed routes are publicly deployed at nikalll.onrender.com.

Built by @taqui.

If this is a team submission, teammates should be credited here with their DEV usernames before publishing.

Top comments (2)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to

Collapse
 
dev_in_the_fog profile image
Jason Y. (dev_in_the_fog) •

Such a clean explanation of a nuanced problem. Bookmarking this for reference, thanks for sharing!