DEV Community

Senanur Çetin
Senanur Çetin

Posted on

My ensemble scored 0.129 on the leaderboard. Here is what explained the gap

My ensemble scored 0.129 on the leaderboard, below the field median. I published it anyway, then tested six explanations for the gap.

Before submitting I wrote down a forecast: 0.143. The graded scores were 0.128 (single model) and 0.129 (ensemble).

Setup

804.5M raw market rows reduced to 292 BigQuery features, walk-forward validation with a one-month embargo, six months held out and read exactly once (+0.15171).

What I found, without spending another submission

  • One hypothesis confirmed: my hold-out sat at the 83rd percentile of period difficulty. De-biasing for that gives +0.14084, within 0.00004 of the walk-forward mean computed a different way.
  • Four falsified, including the two I liked most. One unsettled.
  • Period difficulty explains 46% of the gap. The rest is documented, not explained away.

Five recorded forecasts, five overshoots, all in the same direction. That points to one cause, not five mistakes.

The rule I kept

An internal gain smaller than the fold-to-fold noise tells you which model to prefer, not what a leaderboard will show.

Links

Top comments (2)

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

Writing the forecast down before seeing the graded score is the habit most leaderboard posts skip, and it makes the 46% explanation credible. The 83rd percentile difficulty finding is a good reminder that a single held-out window can mislead more than the model itself. Did you consider scoring across several rolling hold-out windows to see how wide the difficulty spread really is?

iin1006h02

Collapse
 
senanurcetin profile image
Senanur Çetin •

Thanks, and yes, that is essentially where the 83rd percentile figure comes from. The walk-forward folds already act as rolling windows, so I ranked the hold-out period against the difficulty distribution of the other periods instead of trusting it on its own. De-biasing for that brought the hold-out score to +0.14084, within 0.00004 of the walk-forward mean computed a different way, which is what convinced me the window was the outlier, not the model.

What I did not do is re-read the final six-month hold-out in multiple slices. I wanted it read exactly once, so the spread estimate comes from the validation folds only. A proper rolling hold-out scheme is a fair next step, and I would expect it to show the spread is wider than a single window suggests.