We Built an AI to Hunt Earth-Like Planets — Here's How
Finding planets around other stars is hard. Kepler gives us raw light curves — brightness measurements over time — and buried inside that noisy data are tiny dips caused by planets
crossing their star. We built Astrobit 1.0 to find them automatically.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
The Problem
Kepler's Simple Aperture Photometry (SAP) flux is messy. Instrumental systematics, cosmic rays, and quarter-boundary artifacts all look like signals. A naive threshold approach
misses real planets and flags false positives constantly.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Architecture
Raw SAP Flux
↓
[Cleaner] — mask bad cadences + sigma-clip outliers
↓
[Detrender] — per-quarter Savitzky-Golay filter
↓
[BLS Search] — 50k coarse grid → fine refinement → alias check
↓
[Feature Extractor] — SDE, depth, SNR, odd/even, secondary eclipse
↓
[Random Forest Classifier] — trained on 269 labelled stars
↓
[Platt Scaler] — calibrates scores to probabilities
↓
[Vetter] — secondary eclipse, odd/even, recurrence, systematics
↓
Ranked Candidates → submission.csv
Each stage is independent and cacheable — BLS results are cached to CSV so you can interrupt and resume without recomputing.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
What We Built
- Cleaning — mask bad cadences, sigma-clip outliers from raw SAP flux
- Detrending — per-quarter Savitzky-Golay filter. Recovers ~90% of true transit depth vs ~33% with a running median, and avoids edge artifacts at Kepler's quarterly roll boundaries
- BLS Period Search — 50k log-spaced coarse grid + 600-point fine refinement around each peak. Alias checking at 0.5x, 1x, 2x, 3x catches period harmonics. ~100x cheaper than full- resolution search
- Feature Extraction — SDE, transit depth, SNR, odd/even depth ratio, secondary eclipse depth
- Random Forest Classifier — trained on 269 labelled stars, replaces brittle single-SDE-threshold with a multi-feature decision boundary
- Platt Scaling — calibrates raw model scores to actual probabilities on the dev set
- Vetting — secondary eclipse check, odd/even depth consistency, per-quarter recurrence, known systematic period filtering
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Results
| Stage | Impact |
|---|---|
| SG detrending vs running median | ~90% vs ~33% transit depth recovery |
| Coarse-to-fine BLS | ~100x faster than full-resolution search |
| Alias checking | Catches period harmonics at 0.5x–3x |
| RF classifier vs SDE threshold | Multi-feature boundary, fewer false positives |
| Platt scaling | Calibrated confidence scores on dev set |
| Vetting layer | Filters secondary eclipses, systematics, odd/even inconsistencies |
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Lessons Learned
• Detrending matters more than the classifier. A bad detrend corrupts every downstream feature. We spent more time on the SG filter than the ML model — worth it.
• Cache everything. BLS on a full Kepler star takes time. Caching to CSV saved us hours during iteration.
• Single thresholds break. SDE alone is a terrible classifier. The moment we switched to a multi-feature Random Forest, false positive rate dropped significantly.
• Calibration is underrated. Raw model scores are not probabilities. Platt scaling on the dev set made our confidence scores actually trustworthy for ranking candidates.
• Vetting is not optional. The classifier catches most false positives, but secondary eclipse checks and odd/even consistency are cheap and eliminate a whole class of eclipsing
binary contamination.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Stack
Python · scikit-learn · lightkurve · scipy · numpy
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Repo
Top comments (0)