DEV Community

Mystique Racing
Mystique Racing

Posted on Originally published at mystique-racing.com

I Run a Lottery Model That Publishes Its Own Failures — 51 Draws of Mark Six Data

Most lottery prediction sites show you their wins. Mine shows you the misses — every single one, on a public status page. Here's why I built it that way, and what 51 draws of Hong Kong Mark Six data actually look like under an honest model.

The setup

I maintain three independent number generators for Mark Six (6 numbers out of 49, plus an extra number):

  • Science — frequency, recency, and omission-gap weighting over the last 51 draws
  • Physics — ball-machine simulation priors (order statistics, positional bias)
  • Mystic — a deliberately non-statistical baseline (numerology rules) as a control group

Yes, the third one is a control group. If a "mystic" generator ever matches the statistical models' hit rate, that's evidence the statistical models carry no signal at all. So far the mystic baseline is losing, which is the only thing keeping the other two honest.

Last draw's post-mortem (draw 26104)

Winning numbers: [4, 28, 31, 44, 47, 48] + 19

Model Hits Note
Science 0/6 —
Physics 1/6 caught #44
Mystic 0/6 control group

Total: 1 hit vs. 2.2 expected for random picks of 18 numbers out of 49. Lift: −0.55. Below random.

I publish this number anyway. A model you can't audit is just marketing.

What self-calibration looks like

The racing side of the project (HKJC odds modeling) uses EMA-based parameter drift: every signal the model emits gets scored against actual results, and the weights update automatically. Over the last 24 recorded parameter updates, the "late steam" weight oscillated 0.947 → 0.992 → 0.981 while the model tried to correct for a day where 3 of 4 late market movers lost.

Nobody touched those numbers. The model graded its own homework and adjusted.

Draw 26105 (tonight's picks, published in advance)

  • Science: [7, 11, 13, 38, 39, 48] + 43
  • Physics: [7, 27, 30, 34, 44, 48] + 21
  • Mystic: [10, 23, 28, 36, 43, 48] + 3

Consensus across models: 48 (all three), 7 (two of three).

The result and the hit count will be public tomorrow, win or lose.

Why honesty is the actual product

Prediction content is a market for lemons — everyone claims 80% accuracy because nobody audits. The whole project (mystique-racing.com) is built around the opposite bet: full prediction history, full post-mortems, a public /status/ page with pipeline health, and calibration stats that include the losing streaks.

If the model is only as good as random over 200 draws, the site will say so. That's the deal.


Code side: Cloudflare Workers + Pages + KV/D1, cron snapshots every 2 minutes on race days, EMA drift loop in Python. Happy to answer architecture questions in the comments.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

Publishing the misses, and keeping a deliberately meaningless "Mystic" generator as a control, is exactly the right instinct. The control is the most important part of the setup.

Some numbers on how long it takes before the status page can say anything. One 6-number ticket against a 6-of-49 draw has a mean of 0.735 hits and a variance of 0.578, and it gets zero hits 43.6% of the time for any ticket, since every ticket is equally random. So a single draw's lift (-0.55 here) is almost pure noise. To detect a model that genuinely beats random by 20% (0.88 hits per ticket instead of 0.735), you need about 210 draws per model for an 80% chance of noticing it at the 5% level; for a 10% edge, about 840 draws. At two draws a week that's two and eight years, which is worth stating on the status page so readers know when a verdict becomes possible.

Two more things will keep the comparison honest:

  • With three models, the best of the three will beat random more often than any one model would by chance. Compare each model with Mystic on the same draws, and treat "best model this month" as a selection, not a result.
  • The frequency and omission-gap weighting runs on 51 draws, i.e. 306 balls, so each number is expected about 6.2 times with a standard deviation of about 2.5. A number seen 10 times or 3 times is well within chance, so those weights are mostly fitting noise. A chi-square test of the 49 counts against uniform would show whether there is any frequency signal to exploit at all.