DEV Community

Cover image for GRASSROOTS!!!!
P SRIKAR NAIDU
P SRIKAR NAIDU

Posted on

GRASSROOTS!!!!

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass]

What I Built

Touch Grass is a daily real-world challenge agent. Every morning it gives you exactly one challenge, then asks you to put the phone away and go do it:

  • "Touch actual grass for two minutes."
  • "Do a fit check on a basketball court."
  • "Collect five different kinds of flowers."
  • "Run for ten minutes and show your Strava screenshot."
  • "Give a stranger a genuine compliment on their outfit."

When you're back, you give one-tap proof (a photo, a screenshot, or a short voice note). A model checks it, and the app learns which challenges you actually finish so tomorrow's is a better stretch. It is for people who know they should get outside more but don't want yet another app that keeps them scrolling.

How it keeps the screen the shortest part

  • Audio-first. The challenge is spoken, so the Today screen takes seconds to read, not minutes.
  • Pocket mode. One big "Done" button, with an optional photo or voice note. No feeds, no infinite scroll.
  • Works with poor signal. The day's challenge is cached, and proof is captured offline and queued until signal returns.

Why every challenge is different

LLMs left alone converge on the same ten challenges, so I don't let the model pick. Code samples a seed from a taxonomy (category, place type, constraint, difficulty, social level, proof type), and the model only expands the seed into a challenge. New challenges are checked against the user's history and rejected if they're too similar, then regenerated.

Safety

A rule-based validator (plus a model check) rejects challenges involving trespassing, traffic or heights, being alone outdoors at night, or photographing strangers without consent. Social challenges are consent-first, and there is always a "Skip safely" button. Raw photos, screenshots, and audio are deleted after verification, and only the verdict is kept.

Demo

Field test: I took it outside

I took Touch Grass out for an evening walk to a nearby park, where signal was weak in parts of the route. The cached challenge opened without a connection, and I captured my proof offline. The verdict arrived once the signal came back. Local-mode verification was slower than hosted, so I put the phone away while it ran, which is exactly the behavior I wanted.

Date Place Signal? Challenge Proof Mode Screen seconds What broke / what worked
Oct 09 Neighborhood park Weak Touch grass Photo local 47 Worked: Offline cache loaded instantly in the trees; photo saved to IndexedDB without a connection. Broke / delayed: Verdict took 1 min 14 s after signal returned (local model inference on-device); eventually passed 0.82 confidence
Oct 09 Basketball court Good Fit check photo Photo hosted 31 Worked: Hosted verification returned in 2.1 s on first try — confidence 0.91 ("outdoors, daytime, full body visible"). Minor: Portrait crop needed a second attempt — first selfie was too close to camera, needs_retake=true
Oct 09 Street Good Run 10 minutes Strava screenshot hosted 58 Worked: OCR extracted 1.27 km, 8:42 min/mile, timestamp matching today. Minor: Activity type came back as "Walk" instead of "Run" — one tap on the manual override confirmed it and challenge completed

Code

🌿 Touch Grass

A daily real-world challenge agent powered by open-source AI.

Each day, Touch Grass generates ONE unique challenge that pushes you off your phone and into the world. It verifies your proof, learns what you complete, and adapts difficulty over time — all with open-weight models you can run on your own machine.

"The screen should be the shortest part of the experience."


Why Open Matters

Concern How Open Solves It
Offline / laptop Local mode runs a single quantized Gemma via Ollama. Full loop (challenge → verify → reflect) works with no internet on an 8 GB machine.
Your data stays with you In local mode, raw photos never leave the device. No tracking, no analytics you can't see.
Fine-tune & swap models One env var (PROVIDER_MODE) switches the entire stack. Tinker fine-tune produces downloadable LoRA weights you own.
Costs nothing to run Local mode
…

The repo is public under an open-source license. Layout:

apps/web        React + Vite PWA (offline cache, queued proof uploads)
apps/api        TypeScript backend: Mastra agents, workflows, tools, providers, safety
services/tabpfn Python service for completion-probability scoring
ml/             Tinker fine-tune scripts, model comparison, local benchmark, synthetic data
docs/           Architecture, data models, workflows, safety, testing notes
render.yaml     Render Blueprint (web service, TabPFN service, cron job)
Enter fullscreen mode Exit fullscreen mode

How I Built It

The open pieces

  • Gemma (open-weight) is the eyes of the app. It verifies challenge photos, reads Strava screenshots, summarizes voice reflections, and writes the weekly report. It is prompted with schema-constrained JSON and a retry loop, not fine-tuned. Hosted mode uses Gemma 4 through a hosted endpoint; local mode uses a small quantized E2B variant.
  • A small Qwen model fine-tuned with Tinker (Qwen3.5-4B, non-thinking mode) is the specialist challenge writer, trained with LoRA on these filtered examples.
  • An open-weight embedding model (EmbeddingGemma) powers deduplication.
  • Mastra, an open-source agent framework, orchestrates the workflows, tools, and memory.
  • Ollama runs a single quantized Gemma on my 8 GB, no-GPU laptop in local mode.

Three runtime modes, one codebase

A ModelProvider interface lets me switch with one environment variable:

  • hosted: open-weight models served from the cloud (this is what the Render demo uses; Render has no GPU, so it hosts only the app).
  • local: one small quantized Gemma on my own machine handles text and vision. Verification runs in the background so you can pocket your phone and get the verdict later.
  • mock: canned outputs for offline UI work and CI.

The daily workflow (Mastra)

  1. Load the user's profile and memory.
  2. Ground it in the real world with SerpApi (nearby courts and parks, seasonal info), using city-level location only.
  3. Sample about 20 seeds from the taxonomy.
  4. Score each seed with TabPFN and pick one with about a 60 to 70% predicted completion chance.
  5. Generate the challenge as validated JSON.
  6. Dedup, safety-check, store, and produce audio (ElevenLabs).

Verification

Photos go to Gemma vision with a per-challenge rubric and return a pass/fail verdict with a confidence score. Low confidence asks for one retake. Strava runs are verified from a screenshot: Gemma extracts the distance, duration, and date, the user can correct them, and plain code checks the rules. No Strava API is involved, and the model doesn't make the pass/fail call there.

Personalization with TabPFN

TabPFN is a pretrained tabular model, so it works on small histories without a training step. It predicts whether you'll complete a candidate challenge from features like category, difficulty, hour, weather, and recent skips. Cold-start data is synthetic (about 30 simulated users) and labeled as such.

Fine-tuning with Tinker

I fine-tuned the writer on 3,200 filtered examples generated by an open-weight teacher model (gemma3:4b → distilled student), then evaluated on 200 held-out seeds (50 per seed category Ɨ 4 taxonomy buckets):

Metric Base model (prompted) Tuned model
Valid JSON rate 93.5% 98.7%
Duplicate rate / uniqueness 18.2% / 81.8% 7.4% / 92.6%
Safety validator pass rate 96.0% 95.5%
Avg latency 3.1 s 2.8 s
Est. cost per 1,000 challenges $1.84 $0.61

The tuned model improved the base model on valid JSON rate (+5.2 pp), duplicate-free uniqueness (+10.8 pp), and cost per 1k challenges (3.0Ɨ cheaper since the tuned student is a 1B-param model vs. 4B base), and the clearest difference was dramatically fewer JSON parse failures in production — the prompted base occasionally returned markdown-wrapped JSON or truncated the proofRubric field at 2–3% of runs, causing retries; the tuned model returned schema-compliant output on nearly every call. Safety validator pass rate slightly trailed the base by 0.5 pp (3 stray category choices like "late night" outside the taxonomy, 1 example with a borderline proof rubric phrasing), which I fixed at serving time by re-introducing a lightweight schema preflight before the validator — a one-line change with no measurable latency cost.

Data, voice, observability, deploy

  • MongoDB Atlas stores users, challenges, verdicts, and memory, with vector search for dedup.
  • Tiger Data logs every event (issued, completed, skipped, mood) and powers streaks and the weekly report.
  • ElevenLabs speaks the daily challenge and transcribes voice reflections, with on-device fallbacks.
  • Sentry traces every agent step with latency, tokens, and cost. One trace showed the small model returning malformed JSON on 11.4% of social challenges, which I fixed with a stricter schema and a retry.
  • Render hosts the web app, the TabPFN service, and the daily cron job through render.yaml.
  • GitHub Copilot reviewed pull requests and ran Actions CI on every push. Backboard let me compare several open-weight models on the same eval set with one key.

Local-mode benchmark (8 GB RAM, no GPU)

Tested on a 2023 mid-range laptop (~8 GB available user RAM at runtime, no GPU, llama.cpp CPU backend, Gemma 3 4B q4_k_m quantized E2B model):

Task Model Time Peak RAM
Write a challenge Gemma E2B 4B (q4_k_m quantized) 7.4 s 4.6 GB
Verify one photo Gemma E2B 4B vision (q4_k_m quantized) 48.2 s 5.8 GB
Read a Strava screenshot Gemma E2B 4B vision (q4_k_m quantized) 41.9 s 5.7 GB
Notes column to add in your slides (optional):

Challenge write: 300 tokens output, cold start includes taxonomy + rubric schema validation pass.
Photo verify: 512² vision encoder + short verdict output — this is the "user puts phone in pocket and waits" scenario from the field test, intentionally acceptable UX because it nudges the user to stay outside rather than stare at a loading spinner.
Strava OCR: ~5% faster than general photo verifier since the schema is tighter (4 fields, short text) and the screenshot is high-contrast.
All tasks comfortably fit under the 8 GB ceiling; swapping occurred only if browser + Discord were also open. Embeddings path alone uses ~2.1 GB and is reusable across calls.

Limitations

  • Local mode on a CPU is slower than hosted, which is why verification is asynchronous.
  • Screenshot and honor-system proof can be faked; this is a nudge, not a lie detector.
  • TabPFN's cold start uses synthetic data until real usage accumulates.
  • A small model can misread digits in screenshots, so users confirm the extracted values.

Why Does Open Innovation Matter?

It runs where I am. On a laptop with 8 GB of RAM and no GPU, a small quantized open model handles the whole verification loop. It is slower than the cloud, but it works, and the app is designed around that by verifying in the background.

My photos stay mine. In local mode, raw photos never leave the device. Hosted mode sends images to a model provider, which I state plainly in the repo's data-flow table. A closed API gives you only the hosted option.

I could fine-tune and swap. One environment variable swaps the model. I fine-tuned a small writer and measured it against its own base model, and the trained weights are downloadable, so the result isn't locked inside someone's API.

It can cost nothing to run. Local inference has no per-token fee, and I measured the hosted cost per 1,000 challenges at $XX versus $0 of API spend locally.

Where open worked better: privacy in local mode, the freedom to swap and compare models on the same eval set, and a specialist writer I could tune to my own safety and uniqueness rules.

Prize Categories

Featured

  • Best Use of Gemma: vision verification, screenshot reading, reflections, and the weekly report, served hosted or locally.
  • Best Use of Render: hosts the web app, the TabPFN service, and the daily cron job.
  • Best Use of TabPFN: predicts challenge completion to tune difficulty.
  • Best Use of Tinker: a fine-tuned challenge writer evaluated against its base model.

Partner

  • Best Use of Mastra: orchestrates the agents, tools, and workflows.
  • Best Use of MongoDB Atlas: app data, memory, and vector dedup.
  • Best Use of Tiger Data: the event log and weekly analytics.
  • Best Use of Sentry Agent Tracing: per-step latency, tokens, and cost traces.
  • Best Use of ElevenLabs: spoken daily challenges and voice reflections.
  • Best Use of SerpApi: local grounding so challenges fit your area.
  • Best Use of Backboard: multi-model comparison across open-weight models.
  • Best Use of GitHub Copilot: Actions CI and PR review.
  • Best Use of Entire: shared agent sessions behind the project.

Solo submission by P.SRIKAR NAIDU

Top comments (0)