DEV Community

Cover image for I Tracked Weather Forecasts for 3 Weeks to See How Wrong They Really Are
Abdullah Bin Masood
Abdullah Bin Masood

Posted on

I Tracked Weather Forecasts for 3 Weeks to See How Wrong They Really Are

The idea

Weather apps show a 7-day forecast like it's a fact. But how accurate is day 7, really, compared to tomorrow's forecast? I wanted real numbers, not a guess, so I built a small pipeline to find out — and to practice the kind of end-to-end data work I want to do professionally.

The plan: collect daily weather forecasts for several cities, wait for the real weather to happen, compare the two, and see what patterns show up.

How it works

Open-Meteo API --> GitHub Actions (daily) --> Neon Postgres --> Python analysis --> Streamlit dashboard

Every day, a GitHub Actions workflow runs on a schedule and does two things:

  1. Pulls a fresh 8-day forecast (today + 7 days ahead) for three cities — Islamabad, London, and New York — and saves it to a Postgres database (Neon).
  2. Checks which older forecasts now have real weather to compare against, using Open-Meteo's historical archive, and saves those actual values too.

A SQL view joins the two, so every forecast is automatically paired with what actually happened on that date, once it's available. (The "actual" values are ERA5 reanalysis data — a model blended with real observations — not a station reading, which is worth being upfront about.)

Everything runs on free tiers: Open-Meteo needs no API key, Neon's free Postgres is plenty for this volume (about 24 rows a day), and GitHub Actions' free scheduler runs the collector without a server.

What I found, after 3 weeks and 324 scored forecasts

1. Accuracy drops the further out you look — but not evenly.

Same-day forecasts were off by 0.58°C on average. By 7 days out, that grew to roughly 1.5–2.0°C. Expected, but it was useful to see exactly how much worse, and that the growth wasn't perfectly linear — day 7 actually came in slightly better than day 6 in this sample, which (with only 30 rows at that lead time) looks more like noise than a real effect.

2. Location matters as much as lead time.

I didn't expect this one to be so stark. New York's forecasts were consistently the least accurate — close to double Islamabad's error at most lead times (3.22°C vs 1.35°C at day 6). A forecast "7 days out" doesn't mean the same thing in every city.

3. Rain predictions degrade too.

The forecast correctly called rain-or-no-rain 90% of the time same-day, dropping to around 70% by day 6–7.

4. A simple model beat a naive baseline — but not the ones I expected.

I tried to predict how wrong a forecast would be, using lead time and the forecast's own values as features. I compared a plain average-by-lead-time baseline against Ridge regression, Random Forest, and Gradient Boosting, using a time-based train/test split (no shuffling — this is time series data, so the test set has to be later in time than the training set).

Model Test MAE (°C)
Ridge regression 0.495
Random Forest 0.670
Gradient Boosting 0.676
Baseline 0.692

Ridge won by about 28%. The tree-based models didn't beat the baseline — my guess is 261 training rows just isn't enough data for them to find an edge over a simple linear model. I'd expect that to change as more data accumulates.

What I'd do differently

  • Start collecting more cities from day one. Three is enough to see a pattern, not enough to generalize.
  • Log a few more weather variables (humidity, pressure) — they might explain some of the city-to-city gap.
  • The free hosting has real limits worth knowing before you rely on it: Streamlit Cloud and Neon both sleep when idle, so the first load after inactivity is slow.

Try it

The dashboard is live: forecast-tracker.streamlit.app
Code is on GitHub: github.com/abdullahbinmasood702/forecast-tracker

It's still running and collecting data daily, so the numbers above will keep shifting as more weeks come in.

Top comments (0)