DEV Community

sen yin
sen yin

Posted on Originally published at xgmind.com

Your Prediction Table Is Hiding the Confidence


By XG Mind AI (xgmind.com)

Two previews both call a home win. One gives it a 60% chance; the other, 95%. The home side wins. Both models get a tick in the results table.

That table just threw away something important. The second preview was almost certain; the first left serious room for a draw or a defeat. Over many matches, that difference is testable — but a winner-only leaderboard can never show it.

Direction hit rate measures whether the selected outcome happened. Calibration asks whether the confidence was warranted. If a site publishes probabilities, readers should be able to examine both.

A small example with every assumption visible

Take 100 fictional binary forecasts of "home win" versus "not home win." Exactly 60 home wins occur. Model A assigns 0.60 to every home win; Model B assigns 0.95. Both always select home win, so both finish with 60% direction accuracy.

For this binary event, the Brier score is the mean of (p − y)², where y is 1 for a home win and 0 otherwise. Lower is better.

Model A: forecast probability 0.60, 60 of 100 home wins observed, direction accuracy 60%, binary Brier score 0.2400. Model B: forecast probability 0.95, 60 of 100 observed, 60% accuracy, Brier 0.3625.

For A, the calculation is 0.60 × (0.60 − 1)² + 0.40 × 0.60² = 0.24. For B, it is 0.60 × (0.95 − 1)² + 0.40 × 0.95² = 0.3625.

A's probabilities match the observed frequency in this constructed sample. B is overconfident — same direction accuracy, much worse calibration. The winner-only table cannot show that difference.

One more caveat before you run off with this example: neither model distinguishes easy fixtures from hard ones; every probability is constant. A constant forecast can match the overall frequency while being useless at ranking individual matches.

Use the right version of the score

Football match direction normally has three classes — home win, draw, away win. A full probability forecast must assign each a non-negative probability, with the three summing to one.

One common multiclass Brier convention averages the sum of the squared errors across the three classes. Its range is zero to two. Other conventions rescale the score, so state your convention before comparing numbers.

Log loss is another proper scoring rule. It punishes assigning very little probability to an outcome that actually occurs. If clipping is used to avoid numerical trouble at zero, disclose it — changing the clipping rule can change a reported score.

A reliability diagram needs counts

To inspect calibration, group predictions into probability ranges and compare each group's mean predicted probability with its observed event rate. A group averaging 0.60 should see an observed rate near 0.60, subject to uncertainty.

Publish the number of forecasts in each group. A point based on eight matches should not look as conclusive as one based on eight hundred.

Scikit-learn's probability calibration guide explains reliability diagrams and adds an important caution: a lower Brier loss does not necessarily mean better calibration.

Compare like with like

A fair comparison uses the same held-out fixtures, the same publication horizon and the same result definition. Compare a twelve-hour forecast against a baseline that also had twelve hours of information — not against one updated after the team sheets came out.

Look at the full sample and at sensible subgroups. A pooled reliability plot can hide overconfidence in one league and underconfidence in another.

Calibration cannot be invented after the match

Suppose an archived preview contains only "home win" plus a high-confidence badge. It's tempting to turn that badge into 80% after the event and compute a Brier score.

That creates a brand-new forecast that was never published. Unless the numerical mapping was specified and recorded beforehand, it's not a valid probability-performance record.

This applies to us too: our public results page should not be called a calibration dashboard merely because it displays direction and score outcomes. Probability metrics require the relevant pre-match probability vectors to exist and to have been preserved.

A short review routine

Before accepting any model-comparison chart, check five things: the task, the held-out sample, the score convention, the forecast horizon, and the count behind each calibration point. Then read the errors rather than only the headline winner.

And if a model looks unusually accurate, our data-leakage guide is the next stop. Our methodology documents how we keep that boundary intact.

Originally published at xgmind.com/blog — the AI football match-intelligence lab with a fully public prediction ledger at xgmind.com/performance/.

Top comments (0)