My ensemble scored 0.129 on the leaderboard, below the field median. I published it anyway, then tested six explanations for the gap.
Before submitting I wrote down a forecast: 0.143. The graded scores were 0.128 (single model) and 0.129 (ensemble).
Setup
804.5M raw market rows reduced to 292 BigQuery features, walk-forward validation with a one-month embargo, six months held out and read exactly once (+0.15171).
What I found, without spending another submission
- One hypothesis confirmed: my hold-out sat at the 83rd percentile of period difficulty. De-biasing for that gives +0.14084, within 0.00004 of the walk-forward mean computed a different way.
- Four falsified, including the two I liked most. One unsettled.
- Period difficulty explains 46% of the gap. The rest is documented, not explained away.
Five recorded forecasts, five overshoots, all in the same direction. That points to one cause, not five mistakes.
The rule I kept
An internal gain smaller than the fold-to-fold noise tells you which model to prefer, not what a leaderboard will show.
Links
- Case study: MSCapital market forecasting
- Notebook: My hold-out was a lucky stretch (Kaggle)
Top comments (2)
Writing the forecast down before seeing the graded score is the habit most leaderboard posts skip, and it makes the 46% explanation credible. The 83rd percentile difficulty finding is a good reminder that a single held-out window can mislead more than the model itself. Did you consider scoring across several rolling hold-out windows to see how wide the difficulty spread really is?
iin1006h02
Thanks, and yes, that is essentially where the 83rd percentile figure comes from. The walk-forward folds already act as rolling windows, so I ranked the hold-out period against the difficulty distribution of the other periods instead of trusting it on its own. De-biasing for that brought the hold-out score to +0.14084, within 0.00004 of the walk-forward mean computed a different way, which is what convinced me the window was the outlier, not the model.
What I did not do is re-read the final six-month hold-out in multiple slices. I wanted it read exactly once, so the spread estimate comes from the validation folds only. A proper rolling hold-out scheme is a fair next step, and I would expect it to show the spread is wider than a single window suggests.