Most lottery prediction sites show you their wins. Mine shows you the misses — every single one, on a public status page. Here's why I built it that way, and what 51 draws of Hong Kong Mark Six data actually look like under an honest model.
The setup
I maintain three independent number generators for Mark Six (6 numbers out of 49, plus an extra number):
- Science — frequency, recency, and omission-gap weighting over the last 51 draws
- Physics — ball-machine simulation priors (order statistics, positional bias)
- Mystic — a deliberately non-statistical baseline (numerology rules) as a control group
Yes, the third one is a control group. If a "mystic" generator ever matches the statistical models' hit rate, that's evidence the statistical models carry no signal at all. So far the mystic baseline is losing, which is the only thing keeping the other two honest.
Last draw's post-mortem (draw 26104)
Winning numbers: [4, 28, 31, 44, 47, 48] + 19
| Model | Hits | Note |
|---|---|---|
| Science | 0/6 | — |
| Physics | 1/6 | caught #44 |
| Mystic | 0/6 | control group |
Total: 1 hit vs. 2.2 expected for random picks of 18 numbers out of 49. Lift: −0.55. Below random.
I publish this number anyway. A model you can't audit is just marketing.
What self-calibration looks like
The racing side of the project (HKJC odds modeling) uses EMA-based parameter drift: every signal the model emits gets scored against actual results, and the weights update automatically. Over the last 24 recorded parameter updates, the "late steam" weight oscillated 0.947 → 0.992 → 0.981 while the model tried to correct for a day where 3 of 4 late market movers lost.
Nobody touched those numbers. The model graded its own homework and adjusted.
Draw 26105 (tonight's picks, published in advance)
- Science: [7, 11, 13, 38, 39, 48] + 43
- Physics: [7, 27, 30, 34, 44, 48] + 21
- Mystic: [10, 23, 28, 36, 43, 48] + 3
Consensus across models: 48 (all three), 7 (two of three).
The result and the hit count will be public tomorrow, win or lose.
Why honesty is the actual product
Prediction content is a market for lemons — everyone claims 80% accuracy because nobody audits. The whole project (mystique-racing.com) is built around the opposite bet: full prediction history, full post-mortems, a public /status/ page with pipeline health, and calibration stats that include the losing streaks.
If the model is only as good as random over 200 draws, the site will say so. That's the deal.
Code side: Cloudflare Workers + Pages + KV/D1, cron snapshots every 2 minutes on race days, EMA drift loop in Python. Happy to answer architecture questions in the comments.
Top comments (1)
Publishing the misses, and keeping a deliberately meaningless "Mystic" generator as a control, is exactly the right instinct. The control is the most important part of the setup.
Some numbers on how long it takes before the status page can say anything. One 6-number ticket against a 6-of-49 draw has a mean of 0.735 hits and a variance of 0.578, and it gets zero hits 43.6% of the time for any ticket, since every ticket is equally random. So a single draw's lift (-0.55 here) is almost pure noise. To detect a model that genuinely beats random by 20% (0.88 hits per ticket instead of 0.735), you need about 210 draws per model for an 80% chance of noticing it at the 5% level; for a 10% edge, about 840 draws. At two draws a week that's two and eight years, which is worth stating on the status page so readers know when a verdict becomes possible.
Two more things will keep the comparison honest: