DEV Community

Ishan Naik
Ishan Naik

Posted on Originally published at rain.ishannaik.com

Building a Hyperlocal Rain Nowcaster on a $0 Stack: Benchmarking Open-Meteo Against Airport Radar

A rain forecast for Mumbai can tell you to carry an umbrella while your street stays dry. During the monsoon, a useful next-two-hour prediction needs more local evidence than a global model alone can provide.

I built Mumbai Rain, also called पाऊस, around that problem. I collect forecasts, compare them with observations at Mumbai airport, and train a small correction that visitors run in their browsers.

First, a qualification about the title: I do not ingest airport radar. I use METAR surface-weather reports from station VABB, Chhatrapati Shivaji Maharaj International Airport. Airport observers and instruments report present weather; I convert their precipitation codes into hourly rain labels. Radar measures precipitation across an area. A METAR report describes conditions at a station. That distinction sets the limits of the entire experiment.

Mumbai Rain social preview card from the public repository

The source repository contains the data diary, training code, browser arithmetic, and exported model. You can inspect the method and scoreboard without accepting a headline accuracy claim.

Start with a narrower prediction target

Engineers who build global models such as ECMWF and GFS solve atmospheric dynamics over grids. A forecast grid cell cannot resolve each Mumbai street, coastal exposure, or passing monsoon shower. Asking for a coordinate does not create street-level observational coverage.

For this project, I ask a narrower question: given the forecast inputs available now, how likely is an hourly rain observation at the airport?

I use Open-Meteo as the delivery API for two forecast streams:

  • best_match, the primary forecast feed, including precipitation and relative humidity.
  • ecmwf_ifs025, an explicit ECMWF-IFS precipitation benchmark.

I do not train a replacement for ECMWF. I train logistic regression to correct the relationship between forecast features and an independent binary observation. Visitors receive a local-coordinate forecast with a correction learned at VABB. They should not read that as a measured rain probability for every neighbourhood.

Collect forecasts before the outcome arrives

The hourly collection workflow schedules a GitHub Actions run at minute five:

schedule:
  - cron: "5 * * * *"
Enter fullscreen mode Exit fullscreen mode

GitHub can delay scheduled jobs. I treat the cron as a collection schedule, not a real-time delivery guarantee.

In pipeline/log_snapshot.py, I collect at (19.12, 72.85), record the forecast issue time, and append forecast-hour rows to data/log.csv. On subsequent runs, I fetch the last 24 hours of METAR reports and fill labels for matching hours.

flowchart TD
    Clock[Hourly GitHub Actions collection] --> Forecast[Open-Meteo best_match and ECMWF-IFS]
    Forecast --> Diary[data/log.csv: forecasts with blank labels]
    Airport[VABB METAR present-weather reports] --> Labels[Convert UTC to IST; label RA and DZ]
    Labels --> Diary
    Diary --> Filter[Keep labelled forward forecasts]
    Daily[Daily GitHub Actions retraining] --> Filter
    Filter --> Train[Fit logistic regression]
    Train --> Gate[Holdout scores and purged walk-forward folds]
    Gate --> Decision{Promotion criteria met?}
    Decision -->|Yes| Export[public/model.json]
    Decision -->|No| Retain[Retain champion or evaluate raw fallback]
    Export --> CDN[Astro assets on CDN]
    CDN --> Browser[Browser loads model]
    Live[Live Open-Meteo forecast] --> Browser
    Browser --> Math[Six features and sigmoid]
    Math --> Verdict[Hourly rain calls and next-two-hour verdict]

I record both timestamps because they answer different questions:

issued_at          Forecast snapshot time, UTC
valid_at           Forecast target hour, Asia/Kolkata
fc_bestmatch_mm    Primary forecast precipitation
fc_ecmwf_mm        ECMWF forecast precipitation
fc_rh_bestmatch    Forecast relative humidity
hour               Target local hour
recent_rain_mm     Recent precipitation feature
observed_raining   Blank until labelled, then 0 or 1
Enter fullscreen mode Exit fullscreen mode

I can then ask what I knew when I issued a forecast, rather than download a later forecast and pretend I had predicted the outcome.

The collector stores the returned day's hourly series, including hours that may precede the issue time. Before training, I filter those rows out. In matured(), I convert issued_at from UTC to IST and require valid_at >= issued_at. I train on forward forecasts because the browser serves forward hours.

Label rain without grading a forecast against itself

In pipeline/labels.py, I read the METAR wxString field:

Report token Label interpretation
RA, including -RA, +RA, SHRA, TSRA Rain
DZ Drizzle
BR, HZ, FG Mist, haze, fog: not a rain label

I convert each report timestamp from UTC to IST and floor it to the local hour. If I find multiple reports in one hour, I label that hour as rainy if any report includes rain or drizzle.

For a missing report, I leave the label blank. I do not equate missing observations with dry weather.

This gives me an independent target, with several qualifications. Airport reports can miss a shower between observations. A binary label gives me no measured rainfall amount. An hourly label loses sub-hour timing. And a dry airport cannot establish that Bandra, Thane, or a visitor's GPS coordinate stayed dry.

I also keep the recent-rain feature separate from the label. Despite the name recent_rain_mm, I compute that feature from Open-Meteo's precipitation series over the three hours before the current hour. I do not obtain a neighbourhood rain-gauge measurement through that field.

Use Git as the database, within its limits

I keep the diary in data/log.csv. The bot commits the changed CSV and health metrics after collection. In each row, I retain the forecast and later add the observed label.

For one collection point and hourly jobs, I get a readable history and a training dataset that contributors can download without credentials. I also avoid provisioning a database just to collect this experiment.

I accept the costs of that choice. The logger reads and rewrites the CSV. Repeated commits increase repository history. Separate collection and training workflows can encounter competing pushes; their individual concurrency groups do not create a shared database transaction. Git suits the current small workload, but I would revisit storage before collecting hundreds of stations.

Make the daily model earn deployment

The retraining workflow schedules a run at 01:30 UTC. It runs the Python tests, trains the candidate, refreshes health metrics, and commits an artifact if the training decision changes it.

In pipeline/train.py, I require at least 200 labelled forward rows. I also require both rainy and dry examples in training and holdout data.

I fit six features in this order:

x = [best_match_mm, ecmwf_mm, relative_humidity,
     sin(2π hour / 24), cos(2π hour / 24), recent_rain_mm]
Enter fullscreen mode Exit fullscreen mode

I encode the hour as sine and cosine so midnight and 23:00 remain neighbours in the feature space. I keep the feature order identical in Python and src/lib/nowcast.js. Swapping humidity and precipitation would produce valid JavaScript and invalid predictions.

For evaluation, I use the Brier score:

$$
\operatorname{Brier} = \frac{1}{N}\sum_{i=1}^{N}(p_i-y_i)^2
$$

A confident wrong prediction incurs a large penalty. A lower score means lower squared probability error on the evaluated observations.

I compare the candidate against the existing logistic champion, a raw rain call, and climatology: the training set's rain frequency. The main holdout contains the last 20% of rows after time sorting.

Purge by time, not row count

One target hour can appear in several forecast snapshots. If I shuffle rows, I can train on one prediction of an hour and test on another prediction of the same observed event.

I add four expanding-window walk-forward folds. I place boundaries between distinct valid_at timestamps and purge six hours before each test window. I skip folds with fewer than 25 training or test rows, or without both label classes.

Time advances to the right

Fold 1: [training] [6h purge] [test]
Fold 2: [     expanded training     ] [6h purge] [test]
Fold 3: [          expanded training          ] [6h purge] [test]
Fold 4: [               expanded training               ] [6h purge] [test]
Enter fullscreen mode Exit fullscreen mode

For each fold, I calculate Brier skill against the raw call:

$$
\operatorname{BSS} = 1 - \frac{\operatorname{Brier}{candidate}}{\operatorname{Brier}{raw}}
$$

I require at least two usable folds, a positive median skill, and at least two positive fold scores. For a perfect raw reference, the implementation reports zero skill rather than divide by zero.

This excerpt expresses the promotion decision using the functions and score variables from the training module:

# Score candidate, champion, raw call, and climatology on the holdout.
# Calculate fold_skills with purged expanding-window evaluation first.
walk_forward_ok = passes_walk_forward(fold_skills)

if walk_forward_ok and passes_gate(
    cand_brier=cand_b,
    champ_brier=champ_b,
    raw_brier=raw_b,
    clim_brier=clim_b,
):
    with open(MODEL_PATH, "w") as f:
        json.dump(candidate, f, indent=2)
Enter fullscreen mode Exit fullscreen mode

Two implementation details matter when interpreting the benchmark:

  1. The automated raw gate uses best_match, not ECMWF alone. The code thresholds the first feature at 0.3 mm/h to construct its raw binary forecast. I collect ECMWF as a separate benchmark and model input, but I cannot describe raw_brier as an ECMWF-only result.
  2. The holdout gate accepts ties. passes_gate() rejects a candidate whose Brier score exceeds a reference score. The walk-forward median must exceed zero, but the holdout comparisons mean “no worse,” rather than a strict improvement against each reference.

I should also avoid promising that future skill can only improve. I evaluate each deployment against the available historical data. Weather regimes change. The code includes a conditional raw-forecast fallback when a champion loses to the raw forecast or climatology and the walk-forward check permits that decision path.

The four purged folds reduce temporal leakage, but the separate 80/20 row split does not use the same timestamp-grouping and purge logic. Repeated target hours can straddle that split. I treat the walk-forward requirement as essential and retain this limitation when reading the holdout scores.

A scorecard you can inspect

The repository's public/model.json, stamped 2026-10-06T07:40:36Z, records:

Artifact field Value
Training rows 6,904
Holdout rows 1,727
Candidate Brier 0.0936
Raw best-match call Brier 0.2131
Climatology Brier 0.1088
Previous champion Brier 0.0943
Walk-forward median BSS 0.6236

The four fold skill values are 0.6555, 0.8399, 0.5918, and 0.5734. These are historical artifact values, not a benchmark I reran for this article. The row counts include repeated target hours from different issue times; they do not represent 8,631 independent weather events.

Export arithmetic, serve it in the browser

I export six coefficients, an intercept, feature names, and evaluation metadata to public/model.json. The payload fits within roughly 2 KB; the current artifact is smaller. Visitors fetch the model as a static asset, then fetch a live Open-Meteo forecast at their selected coordinates.

For each hour, I calculate:

$$
z = b + \sum_{j=1}^{6} w_j x_j,
\qquad p = \frac{1}{1+e^{-z}}
$$

Here is an illustrative input using the coefficients in the dated artifact, not a recorded weather observation:

best_match precipitation    1.0 mm
ECMWF precipitation         0.6 mm
relative humidity          90.0 percent
local hour                 18:00
hour sine                  -1.0
hour cosine                 0.0
recent precipitation        2.0 mm
Enter fullscreen mode Exit fullscreen mode

The arithmetic is:

z = -16.791748
    + 0.300222 × 1.0
    + 0.667161 × 0.6
    + 0.173399 × 90.0
    + (-0.271461) × (-1.0)
    + (-0.900516) × 0.0
    + 0.041041 × 2.0

z ≈ -0.131776
p = 1 / (1 + exp(0.131776)) ≈ 0.4671
Enter fullscreen mode Exit fullscreen mode

The browser calls an hour rainy at p >= 0.5, so this example falls below the rain-call threshold even though both raw precipitation inputs exceed 0.3 mm/h.

The JavaScript implementation clamps the logit to the same range as Python:

const z = model.intercept + model.weights.reduce(
  (sum, weight, i) => sum + weight * features[i],
  0
);
const clamped = Math.max(-60, Math.min(60, z));
const probability = 1 / (1 + Math.exp(-clamped));
Enter fullscreen mode Exit fullscreen mode

I evaluate each forecast hour in the requested window and build the verdict from those hourly calls. When I display precipitation intensity, I use the raw forecast amount. Logistic regression returns a probability, not millimetres of rain. The maximum hourly probability also should not be interpreted as the probability of any rain across the entire window.

Account for the $0 claim

I use an Astro site, static model assets, keyless forecast requests, and GitHub Actions jobs on a public repository. I do not provision a backend ML inference server. For the browser path, visitors perform the six-feature calculation on their own devices.

That removes a network round trip to an inference service. It does not make the application zero-latency: visitors still download assets, fetch Open-Meteo responses, and wait on network conditions. The repository also includes an on-demand /api/nowcast endpoint for API consumers; the browser inference path does not require that endpoint.

The operating stack fits the project's current free-tier usage. I do not treat $0 as an unlimited-capacity guarantee. Forecast-provider terms, rate limits, hosting allowances, and scheduled-job availability still apply, and domain registration sits outside this runtime budget.

You can inspect the pipeline locally with the repository's documented commands:

uv sync
uv run python -m pipeline.health
uv run python -m pipeline.train
Enter fullscreen mode Exit fullscreen mode

The training command applies the promotion decision and can overwrite public/model.json if the candidate qualifies. Use a working copy when experimenting.

For this project, I care about the chain from a forecast captured before an event to an independent observation and a deployment decision I can inspect. I keep that chain in a CSV, a Python module, and a small JSON artifact. Until I collect observations beyond VABB, I will describe Mumbai Rain as a forecast corrector trained at the airport, with neighbourhood forecasts and explicit geographic limits.

Top comments (0)