A rain forecast for Mumbai can tell you to carry an umbrella while your street stays dry. During the monsoon, a useful next-two-hour prediction needs more local evidence than a global model alone can provide.
I built Mumbai Rain, also called पाऊस, around that problem. I collect forecasts, compare them with observations at Mumbai airport, and train a small correction that visitors run in their browsers.
First, a qualification about the title: I do not ingest airport radar. I use METAR surface-weather reports from station VABB, Chhatrapati Shivaji Maharaj International Airport. Airport observers and instruments report present weather; I convert their precipitation codes into hourly rain labels. Radar measures precipitation across an area. A METAR report describes conditions at a station. That distinction sets the limits of the entire experiment.
The source repository contains the data diary, training code, browser arithmetic, and exported model. You can inspect the method and scoreboard without accepting a headline accuracy claim.
Start with a narrower prediction target
Engineers who build global models such as ECMWF and GFS solve atmospheric dynamics over grids. A forecast grid cell cannot resolve each Mumbai street, coastal exposure, or passing monsoon shower. Asking for a coordinate does not create street-level observational coverage.
For this project, I ask a narrower question: given the forecast inputs available now, how likely is an hourly rain observation at the airport?
I use Open-Meteo as the delivery API for two forecast streams:
-
best_match, the primary forecast feed, including precipitation and relative humidity. -
ecmwf_ifs025, an explicit ECMWF-IFS precipitation benchmark.
I do not train a replacement for ECMWF. I train logistic regression to correct the relationship between forecast features and an independent binary observation. Visitors receive a local-coordinate forecast with a correction learned at VABB. They should not read that as a measured rain probability for every neighbourhood.
Collect forecasts before the outcome arrives
The hourly collection workflow schedules a GitHub Actions run at minute five:
schedule:
- cron: "5 * * * *"
GitHub can delay scheduled jobs. I treat the cron as a collection schedule, not a real-time delivery guarantee.
In pipeline/log_snapshot.py, I collect at (19.12, 72.85), record the forecast issue time, and append forecast-hour rows to data/log.csv. On subsequent runs, I fetch the last 24 hours of METAR reports and fill labels for matching hours.
flowchart TD
Clock[Hourly GitHub Actions collection] --> Forecast[Open-Meteo best_match and ECMWF-IFS]
Forecast --> Diary[data/log.csv: forecasts with blank labels]
Airport[VABB METAR present-weather reports] --> Labels[Convert UTC to IST; label RA and DZ]
Labels --> Diary
Diary --> Filter[Keep labelled forward forecasts]
Daily[Daily GitHub Actions retraining] --> Filter
Filter --> Train[Fit logistic regression]
Train --> Gate[Holdout scores and purged walk-forward folds]
Gate --> Decision{Promotion criteria met?}
Decision -->|Yes| Export[public/model.json]
Decision -->|No| Retain[Retain champion or evaluate raw fallback]
Export --> CDN[Astro assets on CDN]
CDN --> Browser[Browser loads model]
Live[Live Open-Meteo forecast] --> Browser
Browser --> Math[Six features and sigmoid]
Math --> Verdict[Hourly rain calls and next-two-hour verdict]
I record both timestamps because they answer different questions:
issued_at Forecast snapshot time, UTC
valid_at Forecast target hour, Asia/Kolkata
fc_bestmatch_mm Primary forecast precipitation
fc_ecmwf_mm ECMWF forecast precipitation
fc_rh_bestmatch Forecast relative humidity
hour Target local hour
recent_rain_mm Recent precipitation feature
observed_raining Blank until labelled, then 0 or 1
I can then ask what I knew when I issued a forecast, rather than download a later forecast and pretend I had predicted the outcome.
The collector stores the returned day's hourly series, including hours that may precede the issue time. Before training, I filter those rows out. In matured(), I convert issued_at from UTC to IST and require valid_at >= issued_at. I train on forward forecasts because the browser serves forward hours.
Label rain without grading a forecast against itself
In pipeline/labels.py, I read the METAR wxString field:
| Report token | Label interpretation |
|---|---|
RA, including -RA, +RA, SHRA, TSRA
|
Rain |
DZ |
Drizzle |
BR, HZ, FG
|
Mist, haze, fog: not a rain label |
I convert each report timestamp from UTC to IST and floor it to the local hour. If I find multiple reports in one hour, I label that hour as rainy if any report includes rain or drizzle.
For a missing report, I leave the label blank. I do not equate missing observations with dry weather.
This gives me an independent target, with several qualifications. Airport reports can miss a shower between observations. A binary label gives me no measured rainfall amount. An hourly label loses sub-hour timing. And a dry airport cannot establish that Bandra, Thane, or a visitor's GPS coordinate stayed dry.
I also keep the recent-rain feature separate from the label. Despite the name recent_rain_mm, I compute that feature from Open-Meteo's precipitation series over the three hours before the current hour. I do not obtain a neighbourhood rain-gauge measurement through that field.
Use Git as the database, within its limits
I keep the diary in data/log.csv. The bot commits the changed CSV and health metrics after collection. In each row, I retain the forecast and later add the observed label.
For one collection point and hourly jobs, I get a readable history and a training dataset that contributors can download without credentials. I also avoid provisioning a database just to collect this experiment.
I accept the costs of that choice. The logger reads and rewrites the CSV. Repeated commits increase repository history. Separate collection and training workflows can encounter competing pushes; their individual concurrency groups do not create a shared database transaction. Git suits the current small workload, but I would revisit storage before collecting hundreds of stations.
Make the daily model earn deployment
The retraining workflow schedules a run at 01:30 UTC. It runs the Python tests, trains the candidate, refreshes health metrics, and commits an artifact if the training decision changes it.
In pipeline/train.py, I require at least 200 labelled forward rows. I also require both rainy and dry examples in training and holdout data.
I fit six features in this order:
x = [best_match_mm, ecmwf_mm, relative_humidity,
sin(2π hour / 24), cos(2π hour / 24), recent_rain_mm]
I encode the hour as sine and cosine so midnight and 23:00 remain neighbours in the feature space. I keep the feature order identical in Python and src/lib/nowcast.js. Swapping humidity and precipitation would produce valid JavaScript and invalid predictions.
For evaluation, I use the Brier score:
$$
\operatorname{Brier} = \frac{1}{N}\sum_{i=1}^{N}(p_i-y_i)^2
$$
A confident wrong prediction incurs a large penalty. A lower score means lower squared probability error on the evaluated observations.
I compare the candidate against the existing logistic champion, a raw rain call, and climatology: the training set's rain frequency. The main holdout contains the last 20% of rows after time sorting.
Purge by time, not row count
One target hour can appear in several forecast snapshots. If I shuffle rows, I can train on one prediction of an hour and test on another prediction of the same observed event.
I add four expanding-window walk-forward folds. I place boundaries between distinct valid_at timestamps and purge six hours before each test window. I skip folds with fewer than 25 training or test rows, or without both label classes.
Time advances to the right
Fold 1: [training] [6h purge] [test]
Fold 2: [ expanded training ] [6h purge] [test]
Fold 3: [ expanded training ] [6h purge] [test]
Fold 4: [ expanded training ] [6h purge] [test]
For each fold, I calculate Brier skill against the raw call:
$$
\operatorname{BSS} = 1 - \frac{\operatorname{Brier}{candidate}}{\operatorname{Brier}{raw}}
$$
I require at least two usable folds, a positive median skill, and at least two positive fold scores. For a perfect raw reference, the implementation reports zero skill rather than divide by zero.
This excerpt expresses the promotion decision using the functions and score variables from the training module:
# Score candidate, champion, raw call, and climatology on the holdout.
# Calculate fold_skills with purged expanding-window evaluation first.
walk_forward_ok = passes_walk_forward(fold_skills)
if walk_forward_ok and passes_gate(
cand_brier=cand_b,
champ_brier=champ_b,
raw_brier=raw_b,
clim_brier=clim_b,
):
with open(MODEL_PATH, "w") as f:
json.dump(candidate, f, indent=2)
Two implementation details matter when interpreting the benchmark:
-
The automated raw gate uses
best_match, not ECMWF alone. The code thresholds the first feature at0.3 mm/hto construct its raw binary forecast. I collect ECMWF as a separate benchmark and model input, but I cannot describeraw_brieras an ECMWF-only result. -
The holdout gate accepts ties.
passes_gate()rejects a candidate whose Brier score exceeds a reference score. The walk-forward median must exceed zero, but the holdout comparisons mean “no worse,” rather than a strict improvement against each reference.
I should also avoid promising that future skill can only improve. I evaluate each deployment against the available historical data. Weather regimes change. The code includes a conditional raw-forecast fallback when a champion loses to the raw forecast or climatology and the walk-forward check permits that decision path.
The four purged folds reduce temporal leakage, but the separate 80/20 row split does not use the same timestamp-grouping and purge logic. Repeated target hours can straddle that split. I treat the walk-forward requirement as essential and retain this limitation when reading the holdout scores.
A scorecard you can inspect
The repository's public/model.json, stamped 2026-10-06T07:40:36Z, records:
| Artifact field | Value |
|---|---|
| Training rows | 6,904 |
| Holdout rows | 1,727 |
| Candidate Brier | 0.0936 |
| Raw best-match call Brier | 0.2131 |
| Climatology Brier | 0.1088 |
| Previous champion Brier | 0.0943 |
| Walk-forward median BSS | 0.6236 |
The four fold skill values are 0.6555, 0.8399, 0.5918, and 0.5734. These are historical artifact values, not a benchmark I reran for this article. The row counts include repeated target hours from different issue times; they do not represent 8,631 independent weather events.
Export arithmetic, serve it in the browser
I export six coefficients, an intercept, feature names, and evaluation metadata to public/model.json. The payload fits within roughly 2 KB; the current artifact is smaller. Visitors fetch the model as a static asset, then fetch a live Open-Meteo forecast at their selected coordinates.
For each hour, I calculate:
$$
z = b + \sum_{j=1}^{6} w_j x_j,
\qquad p = \frac{1}{1+e^{-z}}
$$
Here is an illustrative input using the coefficients in the dated artifact, not a recorded weather observation:
best_match precipitation 1.0 mm
ECMWF precipitation 0.6 mm
relative humidity 90.0 percent
local hour 18:00
hour sine -1.0
hour cosine 0.0
recent precipitation 2.0 mm
The arithmetic is:
z = -16.791748
+ 0.300222 × 1.0
+ 0.667161 × 0.6
+ 0.173399 × 90.0
+ (-0.271461) × (-1.0)
+ (-0.900516) × 0.0
+ 0.041041 × 2.0
z ≈ -0.131776
p = 1 / (1 + exp(0.131776)) ≈ 0.4671
The browser calls an hour rainy at p >= 0.5, so this example falls below the rain-call threshold even though both raw precipitation inputs exceed 0.3 mm/h.
The JavaScript implementation clamps the logit to the same range as Python:
const z = model.intercept + model.weights.reduce(
(sum, weight, i) => sum + weight * features[i],
0
);
const clamped = Math.max(-60, Math.min(60, z));
const probability = 1 / (1 + Math.exp(-clamped));
I evaluate each forecast hour in the requested window and build the verdict from those hourly calls. When I display precipitation intensity, I use the raw forecast amount. Logistic regression returns a probability, not millimetres of rain. The maximum hourly probability also should not be interpreted as the probability of any rain across the entire window.
Account for the $0 claim
I use an Astro site, static model assets, keyless forecast requests, and GitHub Actions jobs on a public repository. I do not provision a backend ML inference server. For the browser path, visitors perform the six-feature calculation on their own devices.
That removes a network round trip to an inference service. It does not make the application zero-latency: visitors still download assets, fetch Open-Meteo responses, and wait on network conditions. The repository also includes an on-demand /api/nowcast endpoint for API consumers; the browser inference path does not require that endpoint.
The operating stack fits the project's current free-tier usage. I do not treat $0 as an unlimited-capacity guarantee. Forecast-provider terms, rate limits, hosting allowances, and scheduled-job availability still apply, and domain registration sits outside this runtime budget.
You can inspect the pipeline locally with the repository's documented commands:
uv sync
uv run python -m pipeline.health
uv run python -m pipeline.train
The training command applies the promotion decision and can overwrite public/model.json if the candidate qualifies. Use a working copy when experimenting.
For this project, I care about the chain from a forecast captured before an event to an independent observation and a deployment decision I can inspect. I keep that chain in a CSV, a Python module, and a small JSON artifact. Until I collect observations beyond VABB, I will describe Mumbai Rain as a forecast corrector trained at the airport, with neighbourhood forecasts and explicit geographic limits.

Top comments (0)