This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
What I Built
golden-hour tells you the best hour to go outside in the next 24, and a small gemma running on your own machine says it in one sentence. a rules scorer picks the hour. the model never does. it only writes the sentence, and if a number in that sentence doesn't match the forecast, the sentence is thrown away and a plain template is used instead.
$ python -m golden_hour --lat 12.97 --lon 77.59 --place Bengaluru --minutes 45
Go tomorrow at 08:00 in Bengaluru: 23C, dry, wind 6 km/h.
[template] score 95/100, UV 0.9, daylight True
(that run was without the model, so it's the template line. with --model pointing at a gemma file, the same facts come out as a sentence like "Please go for a walk tomorrow at 07:00 in Bengaluru with a temperature of 21C and a 9% rain chance.")
i built it because the easiest reason to stay in is not knowing when the weather is actually fine. a phone weather app gives you 24 numbers. i wanted one hour, one line, and one reason.
Why open models matter here
this is the part i actually care about, so i tried to test it instead of just saying it.
- it runs offline once downloaded. the gemma 3 1B file (Q4_K_M) is about 806 MB. i ran it on a box with 2 cpus and about 2 GB of ram. no gpu, no api key, no account.
- only the forecast leaves your machine. your city goes to open-meteo (free, no key). nothing you type goes to a model provider, because the model is a file on your disk.
- being able to look inside is what made the validator possible. with a local model i could check every number it wrote against the facts it was given and reject it. with a hosted chat box i'd be hoping.
How it works
- forecast from open-meteo, 48 hourly points
-
scorer.pyscores each hour out of 100: it loses points under 18C, over 24C, for rain chance, wind over 15 km/h, UV over 5, and 25 for night. the weights are indocs/scoring.md. they are my judgment, not fitted to anything - the best start hour for your walk length wins
- gemma 3 1B gets the facts and writes one sentence
- the validator checks the sentence against the facts (numbers with their units, the day word, and weather words the forecast doesn't support). any mismatch means the template line is used
- a
whyline says what cost points, for example "lost points to rain chance (-11)". if the best score is under 60, it says there is no good window instead of cheering
What i measured
all on one 2-cpu, 2 GB box, gemma 3 1B (Q4_K_M).
the sentences. 20 cities x 3 walk lengths (30, 60, 120 minutes), one run each, 60 runs:
- 3 runs had no good window at all (nairobi, 92% rain chance at the best hour). there the model is skipped and the app says "no good window", because my first version happily wrote "go for a walk" with 92% rain
- in the other 57, gemma's sentence passed validation 56 times. one was rejected and replaced by the template line
- median time to write a sentence: 2.9 s (2.1 to 4.1 s), including generation
- i also tried gemma 3 270M because it loads in 0.6 s. it ignored the task, so i dropped it
the bug in my own validator. my first validator only checked that every number in the sentence appeared somewhere in the facts. that sounds right and it isn't. at a 06:00 walk it accepted "it is 6C" when the temperature was 21C, because 6 was in the facts (as the hour). it also accepted a swapped "27C with a 21% rain chance", the wrong day, and an invented "sunny". so i corrupted correct sentences in six known ways, 400 each, and counted how many the validator flagged:
| corruption | first version | now |
|---|---|---|
| wrong number | 95% | 100% |
| right numbers, wrong labels | 0% | 100% |
| temperature equals the start hour | 0% | 99% |
| wrong day word (today/tomorrow) | 0% | 100% |
| invented "sunny" | 0% | 100% |
| weather word the facts contradict | 0% | 100% |
correct sentences are still accepted (100%). the fix checks that numbers carry the right unit and value, that the day word is right, and that no weather word appears that the forecast doesn't support. these are corruptions i thought of, so it is a floor on what i tested, not proof it can't miss something.
one real rejection from the run: gemma wrote "...tomorrow at 6:00 PM in Sydney, enjoying the 16C temperature and light rain." the rain chance was 2%. the validator caught "light rain" and the template line went out instead. a 12-hour "6:00 PM" is accepted when it is the same moment as 18:00.
the weights. the scorer's weights are my judgment, so i checked how much they matter: scale every weight randomly between 0.7x and 1.3x and re-pick. over 3,000 re-picks the best hour stayed the same 92% of the time, within an hour 95%, and moved more than 2 hours in 3% (mostly sydney and chennai, where two hours score about the same). that says the pick is stable against small changes in my taste. it does not say the weights match yours.
the demo. the web demo re-implements the scorer in javascript. it matches the python scorer to 2e-14 over 2,000 random hours.
What i did not test, and what is weak
- i did not field test it. it was not used on a real walk by anyone yet. the challenge asks for taking it outside, and i can't claim that. the demo and the numbers above are as far as it goes
- the model is gemma 3, not gemma 4
- the sentences are bland. honestly, the template line is just as useful. the model adds tone, not information. what it gives is a place to check that "a small local model can't invent numbers" is testable
- the validator checks facts (numbers with units, the day, a short list of weather words), not tone, and not any claim outside that list
- the scorer's weights are my taste. someone who likes cold rain would disagree
- 20 cities, one run each, is a sanity check and not a benchmark. i did not check forecast accuracy, that is open-meteo's job
- the web demo runs the scorer in your browser. the sentences shown on it are precomputed on my machine, not generated in your browser
Demo
live, no install: https://maybesomeone-arc18.github.io/golden-hour/demo/standalone.html
type a city, pick a walk length (30 min, 60 min, 2 hours), and it picks the hour from the live forecast and shows why. below it are sample sentences from the local gemma, labelled as precomputed, plus the one the validator rejected.
Code
https://github.com/MaybeSomeone-arc18/golden-hour
plain python, one dependency (llama-cpp-python, only if you want the model). 23 tests, python -m pytest tests. the scripts and raw results behind every number above are in eval/ (run_eval.py, mutation.py, sensitivity.py, parity.py).
AI assistance
this was built with ai assistance. the idea, scoring rules and code were written with an ai agent, and the numbers above come from runs of the code in the repo.
Top comments (0)