Almost every year, India’s east coast—especially Odisha, Andhra Pradesh, and West Bengal—takes a hit from a cyclone. While plenty of tools show where a storm is heading, very few tell you how many people actually live inside the impact zone. Early on, that district-level exposure is what responders really need to prioritize resources.
I built Cyclone Impact Forecaster to bridge that gap. It combines a 6-hour intensity forecast with a district-level exposure ranking for the North Indian Ocean, rendered on a 3D satellite globe.
- Live demo: https://cyclone-forecaster-s015.onrender.com (runs on a free Render instance that sleeps when idle, so the first load takes about 45 to 60 seconds)
- Code: https://github.com/Divyansh0208/Cyclone
A simulated storm placed by hand at 17.2 N, 84.5 E (80 kt), taken on a quiet day when no real storm was active.
How the pipeline works
- Ingesting storm state: Live storms in the North Indian Ocean are pulled from NOAA's IBTrACS ACTIVE file (using JTWC fixes). Because the basin isn't always active, I also built modes to replay historical cyclones or drop a custom storm onto the map by hand.
- 6-hour intensity forecast: An XGBoost model predicts the change in wind speed over the next 6 hours based on current position, wind, central pressure, and their recent trends.
- District exposure model: To estimate human impact, I fitted a 2-parameter log-linear model on 61 historical Indian cyclone records from EM-DAT, combined with Census 2011 district populations and areas. It outputs expected people affected along with a 10-90% prediction interval for every district inside the impact radius.
- 3D visualization: Built on MapLibre GL with satellite imagery and terrain elevation. District markers scale with the estimated affected population, and clicking any district flies the camera down to inspect it.
All of this runs strictly on real datasets—IBTrACS (NOAA NCEI) for tracks, EM-DAT (CRED) for the 61 historical impact records, and the 2011 Census of India. I avoided synthetic data completely; even the test suite runs against a real track excerpt from Cyclone Fani.
Benchmarks
| Metric | Model | Naive Baseline |
|---|---|---|
| Intensity MAE (6 h) | 3.04 +/- 0.20 kt | 3.49 kt (persistence / last value) |
| Impact log-MAE | 2.50 +/- 0.41 | 2.69 (historical mean) |
For the intensity model, I used 5-fold cross-validation grouped by storm, beating the persistence baseline across all 5 folds. For the impact model, I ran 5-fold CV with 20 repeats on the log scale.
Notes from working with 61 data points
- Complex models overfit fast on small disaster datasets. When I first threw a RandomForest at the EM-DAT impact data, it performed worse than simply predicting the historical mean. Stepping back to a 2-parameter log-linear model actually generalized. I even ended up fixing the population density coefficient at 1 (letting it fit freely gave about 1.08, so locking it simplified the model with zero performance penalty). The improvement over the baseline is still modest, so I make that explicit in the UI.
- Random train/test splits cheat on track data. Consecutive 6-hour fixes from the same cyclone look nearly identical. If you do a standard random split, adjacent points end up in both train and test sets, inflating your metrics. Grouping cross-validation folds strictly by storm ID is the only honest way to evaluate it.
- Don't hide wide error bars. In the screenshot above, the top-ranked district's 10-90% range spans from roughly 17,000 to 2.2 million people. That looks massive, but that's the reality of training on 61 noisy historical events. Rather than pretending the point estimate is exact, the UI frames the output as a relative ranking across districts and keeps the error band front and center.
- Be transparent about data lag. The upstream IBTrACS ACTIVE file only rebuilds about three times a week, meaning "live" data can lag by up to two days. The backend caches responses for 15 minutes, falls back to stale data if NOAA's upstream fails, and badges every fix in the UI with its exact age.
-
Keep training and serving feature logic in one place. Both the offline training scripts and the live API import from a single
features.pymodule so feature engineering can't silently drift between the two.
Current limitations
- No track or landfall forecasting yet: Every district inside the impact radius currently receives the same predicted storm wind rather than a distance-decayed wind field.
- Outdated census boundaries: Relying on the 2011 Census means the population figures are 15 years old and several newly carved-out districts aren't represented.
- Wind scale mismatch: JTWC reports 1-minute sustained winds, whereas the India Meteorological Department (IMD) uses a 3-minute averaging period, making storm category labels approximate.
- Not an official warning system: This is an experimental decision-support project. IMD is the sole official authority for cyclone warnings in India.
Help wanted
I have 13 open issues on GitHub—several tagged good first issue—covering a proper distance-decay wind field, better calibration for the 10-90% interval, post-2011 district boundaries, regional language support (Hindi, Odia, Telugu, Bengali), and an accessibility audit.
If you work in disaster response, hydrometeorology, or geospatial ML, I’d really appreciate your critique on what would make this genuinely useful in the field.

Top comments (0)