This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.
What I Built
Baahar (बाहर — outside) finds Bengaluru's next safe outdoor hour from air quality, heat and rain, speaks a ~30-second park briefing, then turns the screen off so you actually go.
There is no feed, no streak, no badge. The only metric that matters is whether you stood up.
It is 2am. I open a weather app: 38°C. I open a map: the air is fine. I open a third: stay inside. Fifteen minutes later I have resolved a question that should have taken twenty seconds, and I have not left the desk.
Fifteen minutes is the problem. Not the weather app.
uv run baahar brief
WAIT Next GO hour is 06:00.
time call comfort NAQI feels why
19:00 WAIT 75 118 26/26° Night-time. Air is moderate, but park gates may be
shut.
20:00 WAIT 76 116 26/26° Night-time. Air is moderate, but park gates may be
shut.
06:00 GO 88 63 20/23° NAQI 63, feels like 23 C.
Then a briefing:
WAIT. Hold off on the walk. Indian NAQI is 118, in the moderate band. The chance of rain is 1%. It is night; park gates may be closed. A later forecast must be checked before making another plan. The next hour to recheck is 06:00. Refresh conditions before walking. Breathing discomfort for people with lung, asthma or heart conditions.
That is the live Bengaluru hour I recorded. Pocket Mode stays off until the current hour is GO. The stills below are a GO hour so you can see the near-black screen; the video in Demo is this WAIT hour.
Tap Pocket the phone. Near-black. One instruction. A timer. You cannot scroll.
Open-Meteo gives you US EPA AQI. India has its own index. Baahar recomputes Indian NAQI from raw pollutant concentrations using CPCB 2014 sub-index breakpoints. The overall index is the worst sub-index, not an average.
The caveat is on every result, because hiding it would make the number quietly wrong:
CPCB breakpoints are defined on 24-hour mean concentrations. Open-Meteo publishes hourly values. Baahar's number is an approximation of official NAQI, not official NAQI.
A related bug only tests caught: CPCB expresses CO in mg/m³ and Open-Meteo reports µg/m³. Without a ÷1000, a plausible 2,000 µg/m³ saturates the index at 500.
Pocket Mode can suggest a species from iNaturalist research-grade records within 5 km of Bengaluru for the current month. The instruction stays short. The evidence sits underneath:
Look for a Chocolate Pansy.
iNaturalist research grade: around a dozen recorded within 5 km of Bengaluru this month. A record means someone logged it nearby, not that you will see it.
The briefing writer never sees that cue. An LLM asked to "include this" will happily rewrite "around a dozen records" as "keep an eye out for the pansies." The fix is removing the temptation, not a better prompt.
Demo
Two environments, same binary. Judges can clone and run both.
Live Bengaluru (this laptop, Open-Meteo, no key).
On 2026-10-10 IST the current hour was WAIT. Next GO hour 06:00. Pocket Mode stayed off. That is the product: it will not start a walk after dark, even though 06:00 in the table is GO.
Recorded screen demo (captions on the page): demo-brief-wait.mp4
git clone https://github.com/Vedant817/baahar.git
cd baahar
uv sync --group dev
uv run baahar brief --city Bengaluru
uv run baahar serve # http://127.0.0.1:8000
Recorded fixtures (what CI runs).
BAAHAR_OFFLINE=1 replays committed Open-Meteo samples so the suite does not depend on tonight's air. 734 tests pass offline. The layout audit hits the live UI when GitHub Actions can.
Outdoor walk.
Not done. docs/FIELD_TEST.md is a blank form. The after-walk journal is wired (uv run baahar journal --markdown) and will paste here when a human fills it. The theme's bonus is "take it outside"; the required demo is a working project and this write-up.
Adversarial pass (personas, not a park).
Four scripted users ran the real CLI/UI and found copy that lied: a WAIT badge that said "Go at 06:00.", "Forecast window 06:00." as a sentence, a journal dump that looked like a completed walk. Those are fixed on main. That is product testing. It is not a Cubbon Park walk.
This is not a hosted Render demo. I did not deploy, so I am not claiming that category.
Code
- Repo: github.com/Vedant817/baahar
- First commit: 2026-10-06 (inside the challenge window)
- License: MIT
- Tests pass offline (
uv run pytest) - Every number below:
eval/RESULTS.md, raw JSON ineval/raw/
How I Built It
weather.py ──┐
├──▶ forecast.py ──▶ features.py ──▶ score.py ──▶ brief.py ──▶ Pocket Mode
air.py ──────┘ (join on (28 tabular (shipped (Gemma / (near-black,
naqi.py timestamp) features) ensemble, template) timer)
(CPCB NAQI) TabPFN, or
heuristic)
Three decisions I would defend:
-
The safety rule is code, not a model output.
features.pyowns GO/WAIT/SKIP, with every threshold anchored to CPCB category edges. The model predicts a physical NAQI band. The boring, auditable part stays code. - Safety is asymmetric. A learned prediction that is less strict than the policy is discarded. The model may talk you out of a walk. It may never talk you into bad air.
-
Safety is repaired after generation.
enforce_safety()appends a real NAQI figure, strips "guaranteed safe"-style hedging, and if the plan says SKIP but the text sounds encouraging, throws the text away.
Go / no-go: predict the CPCB band six hours ahead
8,130 hourly rows, Bengaluru, 2025-11-01 → 2026-10-05, Open-Meteo CAMS + ERA5. Chronological holdout, never shuffled: last 20% is 1,626 rows (2026-07-30 → 2026-10-05). 28 features (13 base + 15 lags / diffs / rolling windows).
| model | accuracy | macro-F1 (4 bands) | moderate recall | skip_as_go | n(SKIP) | fit time |
|---|---|---|---|---|---|---|
| majority class | 0.4047 | 0.1441 | 0.0000 | 0.0 | 24 | <0.1 s |
| persistence | 0.3647 | 0.2492 | 0.2746 | 0.0 | 24 | <0.1 s |
| logistic regression | 0.7897 | 0.5361 | 0.4225 | 0.0 | 24 | 2.0 s |
| random forest | 0.8280 | 0.5780 | 0.5141 | 0.0 | 24 | 3.2 s |
| gradient boosting | 0.8567 | 0.6906 | 0.4789 | 0.0 | 24 | 60.1 s |
| lightgbm | 0.8594 | 0.6347 | 0.5000 | 0.0 | 24 | 23.9 s |
| consensus ensemble (shipped) | 0.8617 | 0.6349 | 0.6620 | 0.0 | 24 | 46.4 s |
| TabPFN 9.1.0 (cpu) | 0.8708 | 0.6193 | 0.5845 | 0.0 | 24 | 358 s |
Accuracy and 4-band macro-F1 are 5-seed means (seeds 0–4). Moderate recall is seed 0. poor has n=3. Two CPCB bands (severe, hazardous) have zero holdout rows. Gradient boosting leads 4-band macro-F1 because it caught 1 of those 3 poor rows; on the three bands with real support the ensemble wins (0.8249 vs TabPFN 0.8236 vs GB 0.7874).
TabPFN is the accuracy leader and is not the engine I ship. It loses moderate recall (0.5845 vs 0.6620) — the under-warning direction for outdoor air. It is also ~7.7× slower to fit. The product cares about walking you into moderate air, not about winning raw accuracy.
The +/- seed spreads are reproducibility across fits on the same split, not statistical significance. Ensemble seed sd is 0.0005; binomial SE at n=1,626 is ~0.009.
Briefings: Gemma ran for real, and the template still ships
36 cases, 12/12/12 GO/WAIT/SKIP from archived Bengaluru hours. Judge is gemini-3.5-flash-lite — not a Gemma model grading its own homework.
| check | template | Gemma 4 (gemma-4-31b-it) |
|---|---|---|
| length ≤ 120 words | 1.000 | 1.000 |
| hallucinated park | 0 | 0 |
| cites NAQI | 1.000 | 1.000 |
| forbidden terms | 0 | 0 |
| GO park name (n=12) | 1.000 | 0.833 |
| blind rubric / 10 | 9.83 | 9.53 |
| latency p50 | 3 ms | 51,730 ms |
The local writer scored higher, named the park on every GO case, and was ~17,000× faster. Gemma is an optional upgrade (--writer gemma). Default is the template. That is not the outcome I expected when I wrote the eval.
The first judge run scored both writers ~3.4/10 because the rubric punished a SKIP briefing for not sounding exciting. A briefing that correctly tells you to stay in is well written. Re-judging the same 72 briefings moved template 3.43 → 9.83 and Gemma 3.33 → 9.53. The harness now refuses a rubric that gives every decision the same score.
What I refused to promote
Hosted Qwen LoRAs learned the templates (validation loss 2.88 → 0.018). They still failed a locked synthetic stress gate (missing air-uncertainty caveats). Serving stays the deterministic writer.
A later ozone TCN on Bapuji Nagar station labels learned typical ozone and missed every Very Poor hour on a held-out April cluster (0/14), same as CAMS. Four train Very Poor hours do not teach a 16-hour spike. I did not ship it.
Tinker is reachable at tinker.thinkingmachines.dev. Running it found 29 of 219 fine-tune examples whose decision label contradicted their own briefing text (night hours: band policy said GO, park gates said WAIT). After the dataset fix, the API returned HTTP 402. I am not claiming the Tinker category.
Why Open Innovation Matters
- "Is it safe to walk?" is a tabular question with a published standard. You want an inspectable model and a confusion matrix, not a wellness score that cannot show its mistakes.
-
The safety number is
skip_as_go_rate. 94% accurate and walking you out on the three worst air days is worse than useless. - Fine-tune and swap. Gemma, a local template, or a hosted LoRA — the product does not change. The writer is a component.
- Cost is the enabler. Open-Meteo is keyless. Gemma's free tier needs no card. That is why a solo first-time builder could ship this in the challenge window.
- Privacy by default. City granularity. No GPS history. No account.
- Open used to reduce screen time. The engineering problem was how little screen this can survive on.
Related work on this same week: Clean Air Walk asks Gemma about Delhi NCR air, and golden-hour keeps a small Gemma from inventing numbers. Baahar's bet is the same honesty, with CPCB NAQI as code and Pocket Mode as the product.
Prize Categories
Listing only what actually ran.
Best Use of Gemma
gemma-4-31b-it generated and evaluated every model briefing (36 cases, machine checks + blind rubric). Open-weight model at the core of the briefing path. I still ship the template as default because it won on park-naming, rubric, and latency.
Best Use of TabPFN
tabpfn==9.1.0 ran on the real 1,626-row chronological holdout, 28 features, 5 seeds: 0.8708 accuracy / 0.6193 4-band macro-F1. Licence accepted by a human, token in .env, CPU override documented. I claim the category because it genuinely ran, and I ship the ensemble because moderate recall is the product metric.
Not claiming
| Category | Why |
|---|---|
| Tinker | Dataset built; LoRA blocked by HTTP 402 after a real run found contradictory labels. |
| Render | Not deployed. |
| ElevenLabs | Client implemented, never called. |
| Arduino / DigitalOcean / others | Not used. |
What Baahar is not
- Not a medical device.
- Not a replacement for the official CPCB advisory.
- Outdoor walk not recorded yet. Live Bengaluru scoring and Pocket refusal are.
Data: Weather and air quality by Open-Meteo (CC BY 4.0). Indian NAQI computed with CPCB 2014 breakpoints. Parks from OpenStreetMap (ODbL). Species cues from iNaturalist research-grade records.
Built by Vedant Mahajan.






Top comments (1)
tr.ee/dev-to