DEV Community

Cover image for Baahar: I built an agent whose success metric is leaving the house
Vedant Mahajan
Vedant Mahajan

Posted on

Baahar: I built an agent whose success metric is leaving the house

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.

What I Built

Baahar (बाहर — outside) finds Bengaluru's next safe outdoor hour from air quality, heat and rain, speaks a ~30-second park briefing, then turns the screen off so you actually go.

There is no feed, no streak, no badge. The only metric that matters is whether you stood up.

It is 2am. I open a weather app: 38°C. I open a map: the air is fine. I open a third: stay inside. Fifteen minutes later I have resolved a question that should have taken twenty seconds, and I have not left the desk.

Fifteen minutes is the problem. Not the weather app.

uv run baahar brief
Enter fullscreen mode Exit fullscreen mode
  WAIT  Next GO hour is 06:00.

  time   call  comfort  NAQI   feels  why
  19:00  WAIT       75   118  26/26°  Night-time. Air is moderate, but park gates may be
                                    shut.
  20:00  WAIT       76   116  26/26°  Night-time. Air is moderate, but park gates may be
                                    shut.
  06:00  GO         88    63  20/23°  NAQI 63, feels like 23 C.
Enter fullscreen mode Exit fullscreen mode

Then a briefing:

WAIT. Hold off on the walk. Indian NAQI is 118, in the moderate band. The chance of rain is 1%. It is night; park gates may be closed. A later forecast must be checked before making another plan. The next hour to recheck is 06:00. Refresh conditions before walking. Breathing discomfort for people with lung, asthma or heart conditions.

That is the live Bengaluru hour I recorded. Pocket Mode stays off until the current hour is GO. The stills below are a GO hour so you can see the near-black screen; the video in Demo is this WAIT hour.

Tap Pocket the phone. Near-black. One instruction. A timer. You cannot scroll.

Baahar briefing

Pocket Mode

Open-Meteo gives you US EPA AQI. India has its own index. Baahar recomputes Indian NAQI from raw pollutant concentrations using CPCB 2014 sub-index breakpoints. The overall index is the worst sub-index, not an average.

The caveat is on every result, because hiding it would make the number quietly wrong:

CPCB breakpoints are defined on 24-hour mean concentrations. Open-Meteo publishes hourly values. Baahar's number is an approximation of official NAQI, not official NAQI.

A related bug only tests caught: CPCB expresses CO in mg/m³ and Open-Meteo reports µg/m³. Without a ÷1000, a plausible 2,000 µg/m³ saturates the index at 500.

Pocket Mode can suggest a species from iNaturalist research-grade records within 5 km of Bengaluru for the current month. The instruction stays short. The evidence sits underneath:

Look for a Chocolate Pansy.
iNaturalist research grade: around a dozen recorded within 5 km of Bengaluru this month. A record means someone logged it nearby, not that you will see it.

The briefing writer never sees that cue. An LLM asked to "include this" will happily rewrite "around a dozen records" as "keep an eye out for the pansies." The fix is removing the temptation, not a better prompt.

Seasonal cue with evidence

Demo

Two environments, same binary. Judges can clone and run both.

Live Bengaluru (this laptop, Open-Meteo, no key).
On 2026-10-10 IST the current hour was WAIT. Next GO hour 06:00. Pocket Mode stayed off. That is the product: it will not start a walk after dark, even though 06:00 in the table is GO.

Recorded screen demo (captions on the page): demo-brief-wait.mp4

Live WAIT briefing

Pocket Mode blocked until GO

git clone https://github.com/Vedant817/baahar.git
cd baahar
uv sync --group dev
uv run baahar brief --city Bengaluru
uv run baahar serve    # http://127.0.0.1:8000
Enter fullscreen mode Exit fullscreen mode

Recorded fixtures (what CI runs).
BAAHAR_OFFLINE=1 replays committed Open-Meteo samples so the suite does not depend on tonight's air. 734 tests pass offline. The layout audit hits the live UI when GitHub Actions can.

Outdoor walk.
Not done. docs/FIELD_TEST.md is a blank form. The after-walk journal is wired (uv run baahar journal --markdown) and will paste here when a human fills it. The theme's bonus is "take it outside"; the required demo is a working project and this write-up.

Adversarial pass (personas, not a park).
Four scripted users ran the real CLI/UI and found copy that lied: a WAIT badge that said "Go at 06:00.", "Forecast window 06:00." as a sentence, a journal dump that looked like a completed walk. Those are fixed on main. That is product testing. It is not a Cubbon Park walk.

This is not a hosted Render demo. I did not deploy, so I am not claiming that category.

After-walk journal

Code

How I Built It

weather.py ──┐
             ├──▶ forecast.py ──▶ features.py ──▶ score.py ──▶ brief.py ──▶ Pocket Mode
air.py ──────┘   (join on        (28 tabular   (shipped       (Gemma /      (near-black,
naqi.py            timestamp)      features)     ensemble,     template)     timer)
(CPCB NAQI)                                     TabPFN, or
                                                heuristic)
Enter fullscreen mode Exit fullscreen mode

Three decisions I would defend:

  1. The safety rule is code, not a model output. features.py owns GO/WAIT/SKIP, with every threshold anchored to CPCB category edges. The model predicts a physical NAQI band. The boring, auditable part stays code.
  2. Safety is asymmetric. A learned prediction that is less strict than the policy is discarded. The model may talk you out of a walk. It may never talk you into bad air.
  3. Safety is repaired after generation. enforce_safety() appends a real NAQI figure, strips "guaranteed safe"-style hedging, and if the plan says SKIP but the text sounds encouraging, throws the text away.

Go / no-go: predict the CPCB band six hours ahead

8,130 hourly rows, Bengaluru, 2025-11-01 → 2026-10-05, Open-Meteo CAMS + ERA5. Chronological holdout, never shuffled: last 20% is 1,626 rows (2026-07-30 → 2026-10-05). 28 features (13 base + 15 lags / diffs / rolling windows).

model accuracy macro-F1 (4 bands) moderate recall skip_as_go n(SKIP) fit time
majority class 0.4047 0.1441 0.0000 0.0 24 <0.1 s
persistence 0.3647 0.2492 0.2746 0.0 24 <0.1 s
logistic regression 0.7897 0.5361 0.4225 0.0 24 2.0 s
random forest 0.8280 0.5780 0.5141 0.0 24 3.2 s
gradient boosting 0.8567 0.6906 0.4789 0.0 24 60.1 s
lightgbm 0.8594 0.6347 0.5000 0.0 24 23.9 s
consensus ensemble (shipped) 0.8617 0.6349 0.6620 0.0 24 46.4 s
TabPFN 9.1.0 (cpu) 0.8708 0.6193 0.5845 0.0 24 358 s

Accuracy and 4-band macro-F1 are 5-seed means (seeds 0–4). Moderate recall is seed 0. poor has n=3. Two CPCB bands (severe, hazardous) have zero holdout rows. Gradient boosting leads 4-band macro-F1 because it caught 1 of those 3 poor rows; on the three bands with real support the ensemble wins (0.8249 vs TabPFN 0.8236 vs GB 0.7874).

TabPFN is the accuracy leader and is not the engine I ship. It loses moderate recall (0.5845 vs 0.6620) — the under-warning direction for outdoor air. It is also ~7.7× slower to fit. The product cares about walking you into moderate air, not about winning raw accuracy.

The +/- seed spreads are reproducibility across fits on the same split, not statistical significance. Ensemble seed sd is 0.0005; binomial SE at n=1,626 is ~0.009.

Briefings: Gemma ran for real, and the template still ships

36 cases, 12/12/12 GO/WAIT/SKIP from archived Bengaluru hours. Judge is gemini-3.5-flash-lite — not a Gemma model grading its own homework.

check template Gemma 4 (gemma-4-31b-it)
length ≤ 120 words 1.000 1.000
hallucinated park 0 0
cites NAQI 1.000 1.000
forbidden terms 0 0
GO park name (n=12) 1.000 0.833
blind rubric / 10 9.83 9.53
latency p50 3 ms 51,730 ms

The local writer scored higher, named the park on every GO case, and was ~17,000× faster. Gemma is an optional upgrade (--writer gemma). Default is the template. That is not the outcome I expected when I wrote the eval.

The first judge run scored both writers ~3.4/10 because the rubric punished a SKIP briefing for not sounding exciting. A briefing that correctly tells you to stay in is well written. Re-judging the same 72 briefings moved template 3.43 → 9.83 and Gemma 3.33 → 9.53. The harness now refuses a rubric that gives every decision the same score.

What I refused to promote

Hosted Qwen LoRAs learned the templates (validation loss 2.88 → 0.018). They still failed a locked synthetic stress gate (missing air-uncertainty caveats). Serving stays the deterministic writer.

A later ozone TCN on Bapuji Nagar station labels learned typical ozone and missed every Very Poor hour on a held-out April cluster (0/14), same as CAMS. Four train Very Poor hours do not teach a 16-hour spike. I did not ship it.

Tinker is reachable at tinker.thinkingmachines.dev. Running it found 29 of 219 fine-tune examples whose decision label contradicted their own briefing text (night hours: band policy said GO, park gates said WAIT). After the dataset fix, the API returned HTTP 402. I am not claiming the Tinker category.

Why Open Innovation Matters

  1. "Is it safe to walk?" is a tabular question with a published standard. You want an inspectable model and a confusion matrix, not a wellness score that cannot show its mistakes.
  2. The safety number is skip_as_go_rate. 94% accurate and walking you out on the three worst air days is worse than useless.
  3. Fine-tune and swap. Gemma, a local template, or a hosted LoRA — the product does not change. The writer is a component.
  4. Cost is the enabler. Open-Meteo is keyless. Gemma's free tier needs no card. That is why a solo first-time builder could ship this in the challenge window.
  5. Privacy by default. City granularity. No GPS history. No account.
  6. Open used to reduce screen time. The engineering problem was how little screen this can survive on.

Related work on this same week: Clean Air Walk asks Gemma about Delhi NCR air, and golden-hour keeps a small Gemma from inventing numbers. Baahar's bet is the same honesty, with CPCB NAQI as code and Pocket Mode as the product.

Prize Categories

Listing only what actually ran.

Best Use of Gemma

gemma-4-31b-it generated and evaluated every model briefing (36 cases, machine checks + blind rubric). Open-weight model at the core of the briefing path. I still ship the template as default because it won on park-naming, rubric, and latency.

Best Use of TabPFN

tabpfn==9.1.0 ran on the real 1,626-row chronological holdout, 28 features, 5 seeds: 0.8708 accuracy / 0.6193 4-band macro-F1. Licence accepted by a human, token in .env, CPU override documented. I claim the category because it genuinely ran, and I ship the ensemble because moderate recall is the product metric.

Not claiming

Category Why
Tinker Dataset built; LoRA blocked by HTTP 402 after a real run found contradictory labels.
Render Not deployed.
ElevenLabs Client implemented, never called.
Arduino / DigitalOcean / others Not used.

What Baahar is not

  • Not a medical device.
  • Not a replacement for the official CPCB advisory.
  • Outdoor walk not recorded yet. Live Bengaluru scoring and Pocket refusal are.

Data: Weather and air quality by Open-Meteo (CC BY 4.0). Indian NAQI computed with CPCB 2014 breakpoints. Parks from OpenStreetMap (ODbL). Species cues from iNaturalist research-grade records.

Built by Vedant Mahajan.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to