DEV Community

Cover image for My weather app said the sky would be clear. I taught an open model to check, then went outside
Uptime Architect
Uptime Architect Subscriber

Posted on

My weather app said the sky would be clear. I taught an open model to check, then went outside

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

Every stargazer knows this evening. The app says clear. You find the warm jacket, drive out past the streetlights, let your eyes adjust for twenty minutes… and look up at a flat grey lid.

Cloud cover is the one number that decides whether going outside at night is worth it, and it's the number weather apps get wrong most often. I run systems for a living, and when a monitor keeps paging you wrongly you don't stop going outside. You measure how wrong it is, and you correct for it.

So I built Clear Tonight?: it takes the cloud forecast your app shows, looks up 32 months of that same forecast's track record at the same spot, and answers one question with an open model: is the sky really going to be clear tonight, and when should I go?

In short: over a year of nights in 6 cities, it cuts the forecast's error by 28–41%. With only a week of history, TabPFN v2 is the only 12-feature model that beats the app. And on Saturday I committed a forecast, went outside, and the sky was clear.

What I Built

  • A live page for 6 cities (Seattle, Minneapolis, Chicago, Denver, Austin, London), refreshed every day. For tonight it shows a verdict (Go 21:00–23:00 or Probably not tonight), an hour-by-hour strip of what your app says against the real chance it's clear, and the Moon.
  • A local version for your own backyard: python -m clear_tonight --lat … --lon … downloads your spot's history, trains on your machine in about 20 seconds, and a local Gemma 4 turns the numbers into a short plan ("…about a 97% chance of really clear skies between 20:00 and 05:00. Since the Moon is 0% lit, it's a great night for faint stars.").
  • A sky-photo log: point it at a photo you took outside and Gemma 4 rates the cloud cover, so you slowly build your own ground truth.

Clear Tonight for Minneapolis on Oct 10: Go 20:00 to 05:00, with the hourly strip of app cloud vs real clear chance

The screen part takes about five seconds. Everything else is outside.

Demo

Live: clear-tonight.onrender.com: pick a city to see tonight's verdict, the hour-by-hour odds, and the backtest for that city.

An 80-second walkthrough, from the backyard to the backtest:

Code

GitHub logo pyaroslav / clear-tonight

Will the sky really be clear tonight? Your weather app's cloud forecast, corrected by TabPFN v2 from its own track record. Local Gemma 4 for the plan.

Clear Tonight?

tests

Will the sky really be clear tonight? Your weather app's cloud forecast, corrected by an open model that learned from the forecast's own track record at your spot.

Live: https://clear-tonight.onrender.com (6 cities, refreshed daily)

Demo video (80 s): https://youtu.be/Qw1fkWXfr4o

Cloud cover decides whether a night walk, a meteor shower or a look at the stars is worth getting off the couch for, and it's the forecast variable weather apps get wrong most often. Clear Tonight takes the forecast that was issued about 24 hours ahead, compares 32 months of those forecasts against what actually happened at the same place, and lets TabPFN v2 learn the pattern of its mistakes. Then it answers one question for tonight: go or don't go, and when?

  • Live page (6 cities, refreshed daily): see site/. Deployed as a static site.
  • Your own spot, fully local: python -m clear_tonight --lat .. --lon ..…

How I Built It

The question, as data

Open-Meteo keeps an archive of what weather models predicted, including the forecast issued about 24 hours ahead (previous_day1) and 48 hours ahead (previous_day2), and a separate archive of what happened (ERA5 reanalysis). For each city I pulled every night hour (8 PM to 4 AM) from February 2024, when the day-ahead archive begins, to last week: about 975 nights and 8,700 hours per city.

Each hour becomes a row of 12 things you'd know the evening before:

  • the forecast cloud cover;
  • humidity and the temperature–dewpoint gap;
  • wind and precipitation;
  • what the older forecast said, and how much it changed between runs;
  • the forecast's view of the whole night;
  • the hour, and the season.

The label is simple: was the sky actually clear (cloud cover ≤ 25%)?

Why TabPFN

TabPFN is a foundation model for tables. Instead of training a model from scratch, you hand it your rows and it predicts in a single forward pass, having learned on millions of synthetic datasets how small tables usually behave. That's exactly the shape of this problem: one small table per place, no time to tune anything, and the ground truth that matters most, your own sky, starts with a handful of nights.

I used TabPFN v2 deliberately. It's the version released under an Apache-2.0-based license (Prior Labs License 1.1, which adds a "Built with PriorLabs-TabPFN" attribution). The newer versions are non-commercial and need a login. v2 runs fully locally on a GPU, or on a CPU for a table this size.

Honest backtest

Walk-forward over the last 12 months: for every month, train only on what came before it and predict that month, the same way you'd use it live.

City Brier: weather app Brier: Clear Tonight Error cut Wasted trips (app → us) Missed clear nights (app → us)
Austin 0.214 0.137 36% 27 → 20 36 → 50
Chicago 0.191 0.130 32% 33 → 19 37 → 53
Denver 0.279 0.164 41% 57 → 35 37 → 48
London 0.194 0.140 28% 16 → 10 47 → 65
Minneapolis 0.233 0.148 36% 29 → 21 51 → 60
Seattle 0.185 0.117 37% 26 → 12 31 → 50

(Brier score: lower is better. A wasted trip is a night that said "go", meaning two clear hours in a row, and turned out cloudy.)

The correction cuts the forecast's probability error by 28–41% in every city, and it roughly halves wasted trips in Seattle and Chicago. The cost is real: it's more cautious, so it misses more clear nights than the app. For a 40-minute drive to a dark-sky park, that's the trade I want. If you'd rather never miss a clear night, use a lower threshold.

And the part I didn't expect: on a full 32 months of history, a plain logistic regression does just as well as TabPFN v2. If you have years of data for one place, you don't need a foundation model.

Where TabPFN actually earns its place

The archive gives any spot on Earth 32 months of history, but its "truth" is ERA5: what happened over a grid cell roughly 30 km wide, not over your backyard or your favourite dark-sky pullout. The truth that really matters is your own sky, and that comes from the photo log one night at a time. So the question that decides which model to use is: what if you only have the last K nights?

Nights of history TabPFN v2 Logistic regression Gradient boosting Recalibrated app number (1 feature) Weather app as-is
7 0.214 0.357 0.259 0.199 0.237
14 0.204 0.316 0.277 0.183 0.237
30 0.186 0.260 0.248 0.179 0.237
60 0.176 0.199 0.253 0.179 0.237
120 0.171 0.161 0.237 0.176 0.237
365 0.153 0.149 0.168 0.174 0.237

(6 cities × 6 months, Brier score.)

Brier score by nights of history: TabPFN v2 beats the weather app from the first week; logistic regression and gradient boosting need two to four months

With a week of data, the classic models given the same 12 features are worse than not correcting at all; they overfit. TabPFN v2, with the same 12 features, beats the raw forecast from the very first week, is the best model of all at two months, and is only overtaken by logistic regression after about four months.

In fairness: for the first month, the simplest fix of all wins, which is just recalibrating the app's own cloud number (0.199 at 7 nights). So the honest recipe is: recalibrate on day one, and switch to TabPFN v2 at around two months. The app uses TabPFN from the start, because it's never worse than the raw forecast and it's the only model that gets better as you add the other 11 features.

Gemma 4: words, not numbers

The local version asks Gemma 4 (gemma4:e4b via Ollama) to turn the hour-by-hour odds into a sentence. In my first try, Gemma read the table and told me the worst two hours were the best, and that the app was more pessimistic than my correction, when it was the other way round. Lesson re-learned: code computes the facts (best window, which forecast is more optimistic, Moon phase) and Gemma only words them. Gemma also does one thing code can't: it looks at a phone photo of the sky and estimates the cloud cover, which becomes a ground-truth log. On the night it rated 3 of 4 photos correctly, and the fourth taught me something (below).

Plumbing

A GitHub Action runs every afternoon. It trains TabPFN v2 on a CPU runner (a single ensemble member there; the backtest used the default four on a GPU), writes tonight.json, and commits it. Render serves site/ as a static site, and the page reads the latest tonight.json straight from the repo, so it's fresh without a redeploy. There's no backend and no API key, and the whole thing costs nothing to run.

Taking it outside

On Saturday morning, Oct 10, I ran it for Minneapolis and committed the forecast to the repo before going anywhere, so I couldn't quietly edit it afterwards. It was about as confident as it gets: 96–98% chance of a clear sky every hour from 8 PM to 5 AM, and a new Moon. Gemma's plan: "Go outside tonight for stargazing… Since the Moon is 0% lit, it's a great night for faint stars. Remember to wear warm layers."

(The warm layers were wrong. It was 80 °F in October.)

At about 7:50 PM I went out to the backyard. The sky was clear from edge to edge, and the first stars were already out before full darkness. It was pleasantly warm, just right for relaxing; I barely moved. The crickets were calming, and the neighbourhood was fantastic.

Clear sky with stars, 7:51 PM

The forecast was right, and the weather model's own 15-minute record for the evening agrees: 0% cloud from 5:45 to 8 PM. One clear night under a confident forecast proves little; the backtest above is the real evidence. But it was worth going out, and that's the point of the app.

Gemma rated 3 of 4 photos correctly, and the fourth taught me something. I ran my sky photos through the photo log. Three came back right: 5–10% cloud, clear sky. The fourth was the same clear sky a few seconds later, but the phone's night mode had lifted the twilight and city glow into a flat grey. Gemma called it 95% cloud, no stars, high confidence, and said the same thing on three re-runs.

The same clear sky, which night mode turned grey; Gemma rated it 95% cloud

That's a useful thing to learn before building on the photo log. A small-data model trained on your own sky is only as good as its labels, and one confident wrong label is worth a lot in a table of 30 rows.

So I fixed it that evening. The photo log now checks every rating against the other photos taken within 15 minutes and against the weather model's cloud cover for that hour. If a rating disagrees with both by 50 points or more, it's held back as needs review instead of being trusted. Re-running the same four photos:

sky 1     0% cloud, stars visible (high)     check: ok
sky 2    10% cloud, stars visible (medium)   check: ok
sky 3    95% cloud, no stars (high)          check: needs review: other photos ~5%, weather model 0%
backyard 10% cloud, no stars (medium)        check: ok
Enter fullscreen mode Exit fullscreen mode

It's the same lesson as the narration: let the model describe what it sees, but check it against what you already know.

I also committed a forecast for the night before, Oct 9, and didn't go out. It will be scored against the final ERA5 data when that's published around Oct 14–15. The early estimates for that night disagree with each other, so I'm not calling it either way yet. I'll post the result in the comments.

Limits, and what's next

  • "Truth" is ERA5, a reconstruction over a grid cell about 30 km wide, not a person looking up from your backyard. That's why the photo log exists, but it doesn't train the model yet. Next: once a spot has a few dozen checked photos, fit TabPFN on those. That's exactly the small-table case the learning curve says it's good at.
  • One forecast model. The input is GFS's day-ahead forecast, because that's what the archive keeps for every day since February 2024. Other models might need less correcting.
  • "Clear" means 25% cloud or less, per hour. Serious astrophotographers would want a stricter line, and casual stargazers a looser one. It's one constant in the code.
  • Cautious by design. It misses more clear nights than the app does, in exchange for fewer wasted trips.
  • One night outside proves nothing on its own. The evidence is the 12-month backtest across 6 cities. The walk is the part a backtest can't show: was it worth going out? (Yes.)

Why Does Open Innovation Matter?

  • The forecast's track record is the product, and it's open too. The whole project rests on Open-Meteo publishing not just today's forecast but every past forecast next to what really happened. Closed weather APIs give you a prediction; open data lets you audit it.
  • An open tabular foundation model fits a one-backyard problem. TabPFN v2's weights run on my desktop's GPU (or a free CI runner) with no account and no per-call cost. That's what makes "train on your spot tonight" a 20-second command instead of a product decision.
  • Your location stays yours. The local version sends nothing but coordinates to the weather archive. The model, the plan and your sky photos stay on your machine.
  • I could check the claims myself. Because everything is open, I could measure exactly where TabPFN helps (little data) and where it doesn't (lots of data), and say so. A hosted black box would have given me a nice number and no idea whether to trust it.

Prize Categories

  • Best Use of TabPFN: TabPFN v2 (Apache-2.0-based license) is the model. It's backtested walk-forward against the raw forecast, climatology, logistic regression and gradient boosting across 6 cities, with a learning curve showing where it wins.
  • Best Use of Render: the live page is a free Render static site serving the repo's site/ folder (a render.yaml Blueprint is included). It shows the forecast the GitHub Action makes with TabPFN v2 every day.
  • Best Use of Gemma: Gemma 4 runs locally to turn the forecast into a plan and to rate sky photos for ground truth, plus a cross-check I added on the night that catches the kind of photo it can misread.

Built with PriorLabs-TabPFN. Weather data by Open-Meteo.com (CC BY 4.0).

— Yaro

Top comments (0)