DEV Community

Cover image for Aurora Odds: should you go outside and look up right now?
Kai Chen
Kai Chen

Posted on

Aurora Odds: should you go outside and look up right now?

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

On 10 May 2024 the northern lights reached Florida, northern India and the Canary Islands. Millions of people saw them. Many more only read about it the next morning, because the alerts on their phones said "G5" and "Kp 9" and nothing about their own sky.

That is the gap I wanted to close. Every aurora app I tried answers a forecaster's question (how disturbed is the magnetic field?) when the question people actually ask is a practical one: is it worth putting my shoes on and going outside, here, right now?

Aurora Odds answers exactly that. Pick your place (or let the browser find it) and you get one verdict for your sky, Go outside now, Bring your phone, Not yet: it's still light out, or Stay in, the sky is quiet, backed by four numbers:

  • the chance the aurora is bright enough for your eyes in the next hour, and for a phone camera, which sees it from much further away;
  • whether it is dark enough where you are, and when it will be;
  • how much of your sky is cloud;
  • which way to face, and how high to look: "lower edge about 8° above the horizon, top near 25°, three fists at arm's length".

There is also a red night-vision mode, because the whole point is to use it outside in the dark without wrecking your eyes' adaptation, and a "tell me when to go out" switch: leave the tab open and it chimes once it is dark, mostly clear and the odds for your eyes pass 40%. The screen can stay in your pocket until then.

The forecast comes from TabPFN, an open-weights tabular foundation model, reading the solar wind that is on its way to Earth.

Demo

Live: https://luoy16002-svg.github.io/aurora-odds/

Aurora Odds replaying the night of 10 May 2024 in London

The page is most fun during a storm, and storms are rare, so it has a replay mode that runs the same model on the solar wind of a real night: London, 10 May 2024 and Chicago, 10-11 October 2024. Clouds in a replay come from Open-Meteo's historical archive for that hour.

The May 2024 replay for London: go outside now

Where to look: the expected arch of the aurora over the poleward horizon

The red night-vision mode on a phone

Code

GitHub logo luoy16002-svg / aurora-odds

Should you go outside and look up right now? A live aurora nowcast for your exact sky, made by TabPFN reading the solar wind.

Aurora Odds

Should you go outside and look up right now? Aurora Odds answers that for your exact sky: the chance the aurora is bright enough for your eyes or your phone camera in the next hour, whether it is dark and clear enough and which way to face and how high to look.

Live page: https://luoy16002-svg.github.io/aurora-odds/ · storm replays: London, 10 May 2024, Chicago, 10-11 October 2024

Aurora Odds replaying 10 May 2024 in London

The forecast is made by TabPFN, an open-weights tabular foundation model reading the solar wind measured at the L1 point 1.5 million km upstream. It is not trained on this problem: it reads 2,000 labelled half-hours from 1998 to 2019 in context (8,000 in the evaluation below) and returns a full probability distribution for the coming hour's geomagnetic activity. One distribution answers every latitude.

How it works

  1. Upstream measurements. NOAA's real-time solar wind feed (magnetic field and plasma from the…

How I Built It

The solar wind gives you a head start

Spacecraft parked at the L1 point, 1.5 million kilometres sunward, measure the solar wind every minute. That wind takes 20 to 80 minutes to reach Earth, so part of the next hour is already measured before it arrives. The most important number is Bz, the north-south direction of the wind's magnetic field. When it points south it connects to Earth's field and pours energy in; when it points north, very little gets through, however fast the wind blows.

I wanted the model to see the world exactly as it can be seen at one moment, so every input is built on an arrival-time axis: each sample is shifted by its own travel time, and at time t the model is only allowed to use samples that had already been measured at L1 by t. That gives four windows: what is measured but still on its way, the last 30 minutes, 30 to 90 minutes ago, and 90 to 180 minutes ago. In each I summarise Bz, the Newell coupling function (a standard estimate of how much energy the wind is delivering), speed, density and pressure, plus the season and time of day. The same code builds 28 years of history from NASA's OMNI data and the live features from NOAA's real-time feed, so there is no training/serving skew to debug at 2 a.m. during a storm.

The target is the highest Hp30 in the coming hour. Hp30 is GFZ Potsdam's half-hourly version of the Kp index: same scale, finer timing, and open-ended, so a superstorm is not squashed at 9.

TabPFN reads history instead of training on it

TabPFN is a transformer pretrained on millions of synthetic tabular problems. You don't train it on your data: you hand it a table of labelled examples (the context) and it predicts new rows in a single forward pass, like in-context learning in a language model.

I gave it 8,000 half-hours from 1998 to 2019 as context (the live page uses 2,000; more on that below). Storms are rare (the next hour reaches Hp30 6 only about 1.2% of the time), so a uniform sample shows the model very few of the moments that matter. I tried a stratified context that over-represents storms and then corrects the predicted distribution back to the real storm frequency. On the 2020-2022 validation years the stratified context won (mean Brier skill 0.46 against 0.44 for Hp30 ≥ 5 to 7), so that's what runs.

The part I like most: TabPFN's regressor returns a full probability distribution over the next hour's Hp30, not a single number. So one forward pass answers every latitude at once. Tromsø needs almost nothing, Edinburgh needs about Kp 5 to see it with the naked eye, London about 7.5. Each of those is just P(Hp30 ≥ k) read from the same distribution.

Did it work?

I tested on 2023 to 2026, years the model never saw, which include the solar maximum, the May 2024 superstorm and the October 2024 storm. I compared it with:

  • LightGBM trained on all 370,000 half-hours from 1998 to 2019 (a classifier per threshold, the usual way to do this);
  • persistence: "the next hour will look like the last half-hour", using the real index with zero delay, which no live app actually has.
Brier skill score (AUC) Hp30 ≥ 4 Hp30 ≥ 5 Hp30 ≥ 6 Hp30 ≥ 7 Hp30 ≥ 8
TabPFN, 8,000-row context 0.56 (0.951) 0.54 (0.975) 0.52 (0.987) 0.58 (0.993) 0.53 (0.990)
TabPFN, 2,000-row context (the live page) 0.54 (0.948) 0.52 (0.973) 0.51 (0.986) 0.56 (0.992) 0.50 (0.992)
LightGBM, trained on 370,431 rows 0.56 (0.951) 0.54 (0.974) 0.51 (0.987) 0.46 (0.989) 0.34 (0.982)
Last half-hour's index (oracle) 0.34 (0.790) 0.33 (0.777) 0.34 (0.776) 0.45 (0.811) 0.44 (0.804)

61,420 half-hours from January 2023 to September 2026. Storm half-hours at each level: 9,345, 3,458, 1,201, 485, 166. Brier skill is relative to always forecasting the long-run frequency; 0 means no skill.

For moderate activity, TabPFN and LightGBM are level. For the rare, strong storms, the ones that bring the aurora to Edinburgh, Chicago or London, TabPFN is clearly better: a Brier skill of 0.58 against 0.46 at Hp30 ≥ 7 and 0.53 against 0.34 at Hp30 ≥ 8. LightGBM saw 46 times more rows, but strong storms are a sliver of them and a tree has very little to split on out there. My guess is that TabPFN's prior, learned from millions of synthetic problems, degrades more gracefully in the tail; either way, it's the tail that decides whether someone in London goes outside.

The live page runs the 2,000-row context, so a free 4-core GitHub runner finishes in about a minute. That costs about 0.02 of skill.

It is not magic. On 10 May 2024 the index climbed to about 6 in the early afternoon while the wind measured upstream still looked calm (a weak field of about 3 nT, barely southward). A model that only sees the wind had nothing to react to and sat at Kp 2 to 3 until the storm's shock arrived at 17:10 UTC; after that it was in the right place within half an hour. It also over-forecasts a little at moderate levels: in the reliability chart the curves for Hp30 ≥ 5 and 6 sit slightly below the diagonal. For a tool whose only job is to tell you when to put your shoes on, I'm fine with erring that way.

The model's forecast against what happened during the May 2024 superstorm

Skill by threshold on 2023-2026

Reliability on 2023-2026

From a number to a sky

The rest is geometry, and it's where most aurora apps stop short.

  • Magnetic latitude. The aurora follows Earth's magnetic field, not lines of latitude, and a simple tilted-dipole model puts Europe several degrees too far north. I precomputed corrected (AACGM-v2) magnetic latitude on a 2° grid and interpolate it in the browser. Edinburgh comes out at 53°, Minneapolis at 54°: that is why Minnesota sees the aurora more often than Scotland despite being 11° further south.
  • What you need. The auroral oval's edge sits near 66.5° magnetic latitude when the field is quiet and moves about 2° equatorward per Kp step. Low on the poleward horizon it is visible from a few degrees further away, and a phone's night mode picks it up further still. I checked the margins against what observers report for Tromsø, Edinburgh, Minneapolis, London and the May 2024 storm.
  • How high to look. Light is emitted 110 to 300 km up. From the distance between you and the oval, a few lines of spherical geometry give the elevation of its lower edge and top, and because the oval is a band at fixed magnetic latitude, it shows up as an arch, highest toward the magnetic pole.
  • Dark and clear. The sun's altitude comes from a standard solar position formula (dark enough below −12°), and cloud cover from Open-Meteo.

A GitHub Actions job reruns the model every ten minutes on a CPU runner and publishes the result with the static page.

Why Does Open Innovation Matter?

Three things in this project only work because the pieces are open.

The model can be hammered for free, at exactly the wrong moment. Aurora traffic is the spikiest traffic there is: nobody looks for months, then a storm hits and everyone looks within the same ten minutes. Because TabPFN's weights are open, the model runs once every ten minutes in a free CI job and writes a small JSON file that a CDN serves to everyone. A million visitors cost the same as one. There is no API key in a secrets store, no per-call bill and no rate limit that kicks in during the one night that matters.

I could change what the model returns, not just what I send it. The label-shift correction works on TabPFN's full predictive histogram: I re-weight its bars by how much more often storms appear in the context than in reality. That needs the raw distribution. An endpoint that returns one number would have made the stratified context unusable, and the "one distribution, every latitude" trick impossible.

Every claim in this post can be checked. The solar wind history is NASA's OMNI data, the index is GFZ's Hp30 (CC BY 4.0), the live feed is NOAA's and the clouds are Open-Meteo's. With open weights on top, python pipeline/evaluate.py reproduces the numbers above on a laptop or a single consumer GPU. If you think my context is badly chosen, or that a different set of features would do better, you can try it in an afternoon and show me.

Prize Categories

Best Use of TabPFN: TabPFN is the forecaster. It reads the labelled history in context and its full predictive distribution turns into the chance of seeing the aurora at every latitude.

Top comments (0)