My ensemble scored 0.129 on the leaderboard, below the field median. I published it anyway, then tested six explanations for the gap.
Before submitting I wrote down a forecast: 0.143. The graded scores were 0.128 (single model) and 0.129 (ensemble).
Setup
804.5M raw market rows reduced to 292 BigQuery features, walk-forward validation with a one-month embargo, six months held out and read exactly once (+0.15171).
What I found, without spending another submission
- One hypothesis confirmed: my hold-out sat at the 83rd percentile of period difficulty. De-biasing for that gives +0.14084, within 0.00004 of the walk-forward mean computed a different way.
- Four falsified, including the two I liked most. One unsettled.
- Period difficulty explains 46% of the gap. The rest is documented, not explained away.
Five recorded forecasts, five overshoots, all in the same direction. That points to one cause, not five mistakes.
The rule I kept
An internal gain smaller than the fold-to-fold noise tells you which model to prefer, not what a leaderboard will show.
Links
- Case study: MSCapital market forecasting
- Notebook: My hold-out was a lucky stretch (Kaggle)
Top comments (1)
Writing the forecast down before seeing the graded score is the habit most leaderboard posts skip, and it makes the 46% explanation credible. The 83rd percentile difficulty finding is a good reminder that a single held-out window can mislead more than the model itself. Did you consider scoring across several rolling hold-out windows to see how wide the difficulty spread really is?
iin1006h02