DEV Community

Shreyash Bhor
Shreyash Bhor

Posted on

Dabba AI: Solving the 8 PM "Aaj Khane Mein Kya Banega?" Crisis with Local Vision & Gemma 2

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Every evening around 8:00 PM, my flatmate stands in front of our open refrigerator door for five full minutes, blankly staring at a random collection of half-used vegetables, leftover paneer, and expired hopes.

The universal Indian bachelor dilemma: "Aaj khane mein kya banega?" (What should we cook today?)

Most cooking apps assume you have 25 exotic ingredients or want a generic 40-minute chef recipe. Food delivery apps just want you to give up and order biryani again.

I built Dabba AI for him.

You snap a quick photo of whatever raw ingredients are sitting in your fridge or kitchen counter, pick your dietary preference (Jain, High Protein, Vegan, or under 15 minutes), and hit generate.

In seconds:

Gemma3 runs locally to inspect the photo and extract only the raw ingredients.

Gemma 3 takes those ingredients and acts like an Indian mom/home cook—crafting a quick, realistic 5-step dabba (tiffin) recipe using standard Indian spice-box staples (haldi, jeera, rai, hing, dhania, etc.).

ElevenLabs reads the recipe aloud in the background so you can cook hands-free with oily or turmeric-stained fingers without touching your laptop screen.

No paid cloud vision APIs. No monthly API billing surprises. The vision and recipe intelligence stay right on the laptop via Ollama.

Demo

Code

GitHub logo shreyash-2006 / hacktober_dev_c1

hactober dev challenge 1 repo

Pull the Ollama models:

Bash

ollama pull gemma3

Clone and install dependencies:

Bash

git clone https://github.com/your-username/dabba-ai.git cd dabba-ai pip install streamlit ollama elevenlabs

Launch the Streamlit app:

Bash

export ELEVENLABS_API_KEY="your-api-key" streamlit run app_gemma.py




How I Built It

The stack is intentional: zero bloated cloud microservices for the core AI logic.

   [ Fridge Photo ] 
          │
          ▼
┌───────────────────┐
│     Gemma3 Vision │  (Local via Ollama: Extracts raw ingredients)
└─────────┬─────────┘
          │ comma-separated tokens
          ▼
┌───────────────────┐
│   Gemma 3 Model   │  (Local via Ollama: Indian home-cook persona)
└─────────┬─────────┘
          │ clean, 5-step recipe text
          ▼
┌───────────────────┐
│ ElevenLabs Audio  │  (Hands-free kitchen voice assistant)
└───────────────────┘
Enter fullscreen mode Exit fullscreen mode
  1. Vision without Cloud Sprawl (Gemma3 via Ollama) Most vision apps immediately send user photos to third-party multi-modal APIs. That adds latency and costs money per snapshot.

I used Gemma3 served locally via Ollama. The prompt is intentionally constrained:

Python
`vision_prompt = (
"Identify all the raw food ingredients visible in this image. "
"List them concisely as comma-separated words only. No sentences, no explanations."
)

vision_response = ollama.chat(
model='paligemma',
messages=[{
'role': 'user',
'content': vision_prompt,
'images': [uploaded_file.getvalue()]
}]
)
ingredients = vision_response['message']['content'].strip()`
Passing strict extraction instructions prevents the vision model from hallucinating cooking steps early—it acts strictly as an ingredient spotter.

  1. Recipe Reasoning Grounded in Indian Kitchens (Gemma 3) The parsed ingredients are piped directly into Gemma 3.

Generic LLMs tend to suggest ingredients you don't have: "Add 2 tbsp of sour cream and fresh tarragon." That fails the Indian kitchen test.

I constrained Gemma 3 to act like a practical home cook: it is strictly allowed to use the identified ingredients plus standard spices from a traditional Indian masala dabba (haldi, jeera, rai, hing, red chilli, salt).

Python
`text_prompt = f"""You are a creative Indian home cook. I have these ingredients: {ingredients}.

Create a fast, practical dabba/tiffin recipe using ONLY these ingredients and standard Indian spices (salt, haldi, jeera, rai, hing, dhania powder, red chilli powder).
Dietary constraint: {dietary_need}.`

Rules:

  • Give the recipe a short, appetising name on the first line.
  • Keep it to exactly 5 numbered steps or fewer.
  • Keep each step under 2 sentences.
  • Use simple conversational language.
  • Do NOT add any intro or outro text.""" This consistently generates actionable, no-nonsense meals in 5 steps or less.
  1. Hands-Free Cooking Audio (ElevenLabs) Anyone who has ever cooked with recipe apps knows the phone screen turns off every 30 seconds, and your fingers are covered in oil, flour, or spice paste. Touching the keyboard or screen is annoying.

With the toggle enabled, the recipe text streams directly to ElevenLabs' eleven_multilingual_v2 model, producing natural spoken instructions right from the Streamlit UI.

Why Open-Weight Models Matter Here
If you build an app that runs every day in your home, relying entirely on paid frontier APIs creates recurring friction. You worry about token usage, billing tiers, or whether an endpoint is down when you are just trying to feed yourself before a late-night coding session.

Running Gemma3 and Gemma 3 locally means:

Zero inference cost: We can snap 10 different angles of the pantry shelf and test recipes without spending a cent.

Privacy by default: Photos of personal living spaces and fridges stay entirely local on your machine.

Speed & Autonomy: No rate limits, no waiting in cloud queues when dinner needs to be made.

How to Run It Locally
Since this runs on local weights rather than a hosted cloud server, you can spin it up on your own machine in 3 steps:

Pull the Ollama models:

Bash
ollama pull gemma3
Clone and install dependencies:

Bash
git clone https://github.com/your-username/dabba-ai.git
cd dabba-ai
pip install streamlit ollama elevenlabs

Launch the Streamlit app:

Bash
export ELEVENLABS_API_KEY="your-api-key"
streamlit run app_gemma.py

No more staring into the fridge void. Snap, cook, eat.``

Why it matters?

Open innovation is what turns AI from an expensive subscription product into everyday household plumbing. When you build utility tools for simple daily friction—like deciding what to cook dinner with—relying entirely on closed, metered APIs breaks the whole premise:

  • No Billing Anxiety for Daily Living: If snapping a photo of your vegetable tray costs a fraction of a dollar in token fees, you will instinctively avoid using the tool. With open-weight models like PaliGemma and Gemma 2 via Ollama, inference costs nothing once downloaded. You can snap four different angles of a half-empty crisper drawer without watching an API billing counter tick up.

  • Privacy Inside Personal Living Spaces: Photos of your kitchen counter, fridge shelves, and home surroundings are inherently personal. Open innovation allows multi-modal vision to execute directly on local silicon, ensuring candid photos never get transmitted to third-party cloud servers or stored in model training logs.

  • Resilience and True Offline Autonomy: Cooking happens when you are hungry, not when an upstream cloud status page is operational. Open weights give developers complete control over execution, latency, and system dependencies, so the tool stays dependable regardless of network drops or API deprecations.

Prize Categories

  • Best Gemma Application: For designing a multi-step pipeline powered locally by Google's open models—using Gemma3 for multi-modal ingredient extraction and Gemma 3 for constrained culinary reasoning.

  • Best Audio / Speech Integration: For using ElevenLabs (eleven_multilingual_v2) to turn structured recipes into hands-free spoken audio so the cook never has to smudge a screen with spice-covered hands.

Top comments (1)

Collapse
 
respect17 profile image
Kudzai Murimi •

I built almost the exact same idea this weekend, a kitchen assistant on local Gemma. Didn't think of using vision to actually look at the fridge though, that's a better version of the problem than just typing ingredients.