DEV Community

Mystique Racing
Mystique Racing

Posted on Originally published at mystique-racing.com

Building a cold-horse radar: steam moves, model divergence, and honest calibration

Research & education only — not betting advice. 18+.

Most "prediction" content in horse racing is storytelling. We wanted a system that answers a narrow, testable question instead:

When our statistical model disagrees with the market about a horse, and the market then moves toward our model's view before the race — can we detect that in near real time?

That detection layer is what we call the cold-horse radar (冷門雷達). This post is the engineering breakdown.

Signal 1: steam moves (落飛)

A steam move is a significant odds drop in a short window — e.g. a horse opening at 29.0 and steaming to 19.0 (a -34.5% move). The implied probability moves from ~3.4% to ~5.3%.

Implementation notes:

  • Poll the public odds board at fixed intervals; snapshot everything
  • Compute delta = (odds_now - odds_open) / odds_open
  • Threshold: flag at |delta| >= 20% — below that is noise in the pools we watch
  • Direction matters: a drop (落飛) means money arrived; a drift (升飛) means the market is cooling

Signal 2: model-vs-market divergence

Our model (gradient boosting over ratings, speed figures, going, draw, weight, jockey/trainer features) outputs a calibrated win probability. The market's implied probability comes from overround-removed odds.

divergence = model_probability - market_implied_probability
Enter fullscreen mode Exit fullscreen mode

We flag when divergence >= 6 percentage points and implied probability <= 22% (roughly 7/1 or longer). The 6pp threshold was tuned down from 8pp after backtesting showed the stricter filter starved the radar of signals on ordinary race days.

Why calibration is the whole game

A divergence signal is only meaningful if your model's "22%" actually means 22%. We verify with:

  • Brier score per meeting and per season (public on our blog)
  • Calibration curves — predicted probability buckets vs realized frequency
  • Time-series split validation — always train on the past, validate on the future

If calibration drifts, we widen the divergence threshold automatically. A miscalibrated radar is worse than no radar — it manufactures false confidence.

What we publish

Every signal, hit or miss, goes into an open scorecard on the site. Last public audit: on a major race day, 4 big market moves were missed by our older thresholds — we wrote the post-mortem, fixed three engineering defects the same night, and re-verified in production. Honest losses are the only kind of track record worth having.

Full methodology series (Cantonese, Japanese and English) lives at mystique-racing.com/blog — newest entry: Ratings, Speed Figures and Market Odds: A Map for Racing Data Science.


Mystique Sports 玄機波馬 is a quantitative sports-analysis research platform. Everything above is reproducible methodology, not tipping.

Top comments (0)