DEV Community

Senanur Çetin
Senanur Çetin

Posted on

My ensemble scored 0.129 on the leaderboard. Here is what explained the gap

My ensemble scored 0.129 on the leaderboard, below the field median. I published it anyway, then tested six explanations for the gap.

Before submitting I wrote down a forecast: 0.143. The graded scores were 0.128 (single model) and 0.129 (ensemble).

Setup

804.5M raw market rows reduced to 292 BigQuery features, walk-forward validation with a one-month embargo, six months held out and read exactly once (+0.15171).

What I found, without spending another submission

  • One hypothesis confirmed: my hold-out sat at the 83rd percentile of period difficulty. De-biasing for that gives +0.14084, within 0.00004 of the walk-forward mean computed a different way.
  • Four falsified, including the two I liked most. One unsettled.
  • Period difficulty explains 46% of the gap. The rest is documented, not explained away.

Five recorded forecasts, five overshoots, all in the same direction. That points to one cause, not five mistakes.

The rule I kept

An internal gain smaller than the fold-to-fold noise tells you which model to prefer, not what a leaderboard will show.

Links

Top comments (1)

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

Writing the forecast down before seeing the graded score is the habit most leaderboard posts skip, and it makes the 46% explanation credible. The 83rd percentile difficulty finding is a good reminder that a single held-out window can mislead more than the model itself. Did you consider scoring across several rolling hold-out windows to see how wide the difficulty spread really is?

iin1006h02