DEV Community

Tanmaya sree Chirra
Tanmaya sree Chirra

Posted on

Finding Exoplanets in Noisy Data with Machine Learning

We Built an AI to Hunt Earth-Like Planets — Here's How

Finding planets around other stars is hard. Kepler gives us raw light curves — brightness measurements over time — and buried inside that noisy data are tiny dips caused by planets
crossing their star. We built Astrobit 1.0 to find them automatically.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

The Problem

Kepler's Simple Aperture Photometry (SAP) flux is messy. Instrumental systematics, cosmic rays, and quarter-boundary artifacts all look like signals. A naive threshold approach
misses real planets and flags false positives constantly.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Architecture

Raw SAP Flux

[Cleaner] — mask bad cadences + sigma-clip outliers

[Detrender] — per-quarter Savitzky-Golay filter

[BLS Search] — 50k coarse grid → fine refinement → alias check

[Feature Extractor] — SDE, depth, SNR, odd/even, secondary eclipse

[Random Forest Classifier] — trained on 269 labelled stars

[Platt Scaler] — calibrates scores to probabilities

[Vetter] — secondary eclipse, odd/even, recurrence, systematics

Ranked Candidates → submission.csv

Each stage is independent and cacheable — BLS results are cached to CSV so you can interrupt and resume without recomputing.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

What We Built

  1. Cleaning — mask bad cadences, sigma-clip outliers from raw SAP flux
  2. Detrending — per-quarter Savitzky-Golay filter. Recovers ~90% of true transit depth vs ~33% with a running median, and avoids edge artifacts at Kepler's quarterly roll boundaries
  3. BLS Period Search — 50k log-spaced coarse grid + 600-point fine refinement around each peak. Alias checking at 0.5x, 1x, 2x, 3x catches period harmonics. ~100x cheaper than full- resolution search
  4. Feature Extraction — SDE, transit depth, SNR, odd/even depth ratio, secondary eclipse depth
  5. Random Forest Classifier — trained on 269 labelled stars, replaces brittle single-SDE-threshold with a multi-feature decision boundary
  6. Platt Scaling — calibrates raw model scores to actual probabilities on the dev set
  7. Vetting — secondary eclipse check, odd/even depth consistency, per-quarter recurrence, known systematic period filtering

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Results

Stage Impact
SG detrending vs running median ~90% vs ~33% transit depth recovery
Coarse-to-fine BLS ~100x faster than full-resolution search
Alias checking Catches period harmonics at 0.5x–3x
RF classifier vs SDE threshold Multi-feature boundary, fewer false positives
Platt scaling Calibrated confidence scores on dev set
Vetting layer Filters secondary eclipses, systematics, odd/even inconsistencies

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Lessons Learned

Detrending matters more than the classifier. A bad detrend corrupts every downstream feature. We spent more time on the SG filter than the ML model — worth it.
Cache everything. BLS on a full Kepler star takes time. Caching to CSV saved us hours during iteration.
Single thresholds break. SDE alone is a terrible classifier. The moment we switched to a multi-feature Random Forest, false positive rate dropped significantly.
Calibration is underrated. Raw model scores are not probabilities. Platt scaling on the dev set made our confidence scores actually trustworthy for ranking candidates.
Vetting is not optional. The classifier catches most false positives, but secondary eclipse checks and odd/even consistency are cheap and eliminate a whole class of eclipsing
binary contamination.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Stack

Python · scikit-learn · lightkurve · scipy · numpy

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Repo

🔗 github.com/25wh1a6678-art/Astrobit_1.0

Top comments (0)