<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sen yin</title>
    <description>The latest articles on DEV Community by sen yin (@sen_yin_caa8aa3b2b41be18f).</description>
    <link>https://dev.to/sen_yin_caa8aa3b2b41be18f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4154792%2Fa8713b32-3536-4525-97eb-0b60f7e06f6f.png</url>
      <title>DEV Community: sen yin</title>
      <link>https://dev.to/sen_yin_caa8aa3b2b41be18f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sen_yin_caa8aa3b2b41be18f"/>
    <language>en</language>
    <item>
      <title>xG Won't Tell You Who Deserved to Win — But It Asks Good Questions</title>
      <dc:creator>sen yin</dc:creator>
      <pubDate>Fri, 02 Oct 2026 04:24:05 +0000</pubDate>
      <link>https://dev.to/sen_yin_caa8aa3b2b41be18f/xg-wont-tell-you-who-deserved-to-win-but-it-asks-good-questions-3h6o</link>
      <guid>https://dev.to/sen_yin_caa8aa3b2b41be18f/xg-wont-tell-you-who-deserved-to-win-but-it-asks-good-questions-3h6o</guid>
      <description>

&lt;p&gt;&lt;em&gt;By XG Mind AI (xgmind.com)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A side grinds out a 1–0 victory, posts 0.7 xG to the opponent's 1.5 — and the comments immediately fill with "they got lucky." That instinct is useful. It just isn't finished thinking yet.&lt;/p&gt;

&lt;p&gt;Maybe the opponent generated most of its xG after going behind and chasing the game. Maybe a penalty explains the whole gap. Maybe the two numbers came from providers that see the same shots very differently. The scoreline and the chance model are answering different questions, and an xG audit should surface those questions — not hand you a verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start by identifying the measurement
&lt;/h2&gt;

&lt;p&gt;Expected goals assigns a scoring probability to a shot that was recorded. Typical features include shot location, angle, body part and the action before the shot. Some models also factor in goalkeeper and defender positions.&lt;/p&gt;

&lt;p&gt;Hudl StatsBomb's xG explainer gives a telling example: a basic model might rate a shot at 0.30, while a richer model that knows the keeper is out of position rates the same shot at 0.65.&lt;/p&gt;

&lt;p&gt;When two totals disagree, resist the urge to average them into manufactured agreement. First establish whether both providers counted the same shots and used the same definitions.&lt;/p&gt;

&lt;p&gt;There's another boundary to respect: xG only covers moves that became shots. A missed cutback, a last-ditch tackle before the shot — those can tell a tactically important story without adding a single decimal to conventional shot-based xG.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use a small audit table, not a verdict
&lt;/h2&gt;

&lt;p&gt;The following is a &lt;strong&gt;fictional five-match example&lt;/strong&gt;. The numbers illustrate interpretation; they are not any club's real results or an XG Mind performance claim.&lt;/p&gt;

&lt;p&gt;Team A: goals scored 10, xG for 6.0, xG against 7.0, xG difference −1.0. Team B: goals scored 4, xG for 7.0, xG against 5.5, xG difference +1.5.&lt;/p&gt;

&lt;p&gt;A scoreboard-only reading favours Team A's attack. The xG totals raise different questions: how did Team A convert those opportunities into ten goals, and why did Team B convert so few?&lt;/p&gt;

&lt;p&gt;Finishing quality, goalkeeping, event-model limitations and plain variation are all on the table — and these totals alone can't separate them. Read the table as a prompt to investigate, not as a forecast that Team B is about to improve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Regression is not a deadline
&lt;/h2&gt;

&lt;p&gt;Regression to the mean is often sold like a debt collector: a team has scored "too many" goals and now must pay them back. Probability doesn't work that way.&lt;/p&gt;

&lt;p&gt;An unusually hot stretch can contain both genuine signal and positive noise. When the observation was picked precisely &lt;em&gt;because&lt;/em&gt; it was extreme, later observations tend to be less extreme. That does not mean the next match reverses the previous one.&lt;/p&gt;

&lt;p&gt;There is no universal six-match threshold, no law that 135% of xG guarantees an imminent collapse, and no guarantee that a goals-to-xG ratio converges to exactly one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three checks before you interpret a gap
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Penalties.&lt;/strong&gt; Keep total xG and non-penalty xG separate when it matters. Don't mechanically subtract some universal constant. And winning penalties can reflect attacking behaviour — they're not automatically "unearned."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Game state.&lt;/strong&gt; A team protecting a lead will happily concede territory and shots. Check &lt;em&gt;when&lt;/em&gt; the chances happened, look for red cards and score changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Opposition and sample.&lt;/strong&gt; Five fixtures against weak opponents are not interchangeable with five against strong ones. Competition, home/away split and time window all matter.&lt;/p&gt;

&lt;p&gt;These are checks, not automatic adjustments. If you claim to have corrected for game state or opposition, explain the procedure instead of just placing "adjusted" in front of the metric.&lt;/p&gt;

&lt;h2&gt;
  
  
  What post-shot xG adds
&lt;/h2&gt;

&lt;p&gt;Post-shot models incorporate information about the shot &lt;em&gt;after&lt;/em&gt; it was struck — placement, sometimes velocity. They're often used to study goalkeeping. But the names PSxG and xGOT don't guarantee identical definitions across providers, so don't assume compatibility.&lt;/p&gt;

&lt;p&gt;One spectacular match is worth describing as one spectacular match. It doesn't establish an inevitable reversal next week.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the forecast and the audit apart
&lt;/h2&gt;

&lt;p&gt;An audit after full-time has information that didn't exist before kickoff. It can explain recorded chance creation, but it can't retrospectively validate a forecast — unless the original forecast and its inputs were preserved.&lt;/p&gt;

&lt;p&gt;The practical routine is short: identify the provider, inspect penalties and game state, check the fixture window, and read the shot pattern. Then ask whether the written conclusion is stronger than the evidence allows.&lt;/p&gt;

&lt;p&gt;That's also why we keep our workings visible. Our &lt;a href="https://xgmind.com/methodology/" rel="noopener noreferrer"&gt;methodology&lt;/a&gt; documents how our forecasts are produced, the &lt;a href="https://xgmind.com/performance/" rel="noopener noreferrer"&gt;performance ledger&lt;/a&gt; records outcomes publicly, and our &lt;a href="https://xgmind.com/blog/missing-football-data-is-not-zero/" rel="noopener noreferrer"&gt;missing-data guide&lt;/a&gt; covers what to do when a comparison lacks one side's metrics. An audit you can't reproduce is just an opinion with a spreadsheet attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;em&gt;Originally published at &lt;a href="https://xgmind.com/blog/xg-the-ultimate-lie-detector-football-audit/" rel="noopener noreferrer"&gt;xgmind.com/blog&lt;/a&gt; — the AI football match-intelligence lab with a fully public prediction ledger at &lt;a href="https://xgmind.com/performance/" rel="noopener noreferrer"&gt;xgmind.com/performance/&lt;/a&gt;.&lt;/em&gt;
&lt;/h2&gt;

</description>
      <category>datascience</category>
      <category>statistics</category>
      <category>football</category>
      <category>analytics</category>
    </item>
    <item>
      <title>Your Prediction Table Is Hiding the Confidence</title>
      <dc:creator>sen yin</dc:creator>
      <pubDate>Fri, 02 Oct 2026 04:18:16 +0000</pubDate>
      <link>https://dev.to/sen_yin_caa8aa3b2b41be18f/your-prediction-table-is-hiding-the-confidence-262l</link>
      <guid>https://dev.to/sen_yin_caa8aa3b2b41be18f/your-prediction-table-is-hiding-the-confidence-262l</guid>
      <description>

&lt;p&gt;&lt;em&gt;By XG Mind AI (xgmind.com)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two previews both call a home win. One gives it a 60% chance; the other, 95%. The home side wins. Both models get a tick in the results table.&lt;/p&gt;

&lt;p&gt;That table just threw away something important. The second preview was almost certain; the first left serious room for a draw or a defeat. Over many matches, that difference is testable — but a winner-only leaderboard can never show it.&lt;/p&gt;

&lt;p&gt;Direction hit rate measures whether the selected outcome happened. Calibration asks whether the &lt;em&gt;confidence&lt;/em&gt; was warranted. If a site publishes probabilities, readers should be able to examine both.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small example with every assumption visible
&lt;/h2&gt;

&lt;p&gt;Take &lt;strong&gt;100 fictional binary forecasts&lt;/strong&gt; of "home win" versus "not home win." Exactly 60 home wins occur. Model A assigns 0.60 to every home win; Model B assigns 0.95. Both always select home win, so both finish with 60% direction accuracy.&lt;/p&gt;

&lt;p&gt;For this binary event, the Brier score is the mean of &lt;code&gt;(p − y)²&lt;/code&gt;, where &lt;code&gt;y&lt;/code&gt; is 1 for a home win and 0 otherwise. Lower is better.&lt;/p&gt;

&lt;p&gt;Model A: forecast probability 0.60, 60 of 100 home wins observed, direction accuracy 60%, binary Brier score 0.2400. Model B: forecast probability 0.95, 60 of 100 observed, 60% accuracy, Brier 0.3625.&lt;/p&gt;

&lt;p&gt;For A, the calculation is &lt;code&gt;0.60 × (0.60 − 1)² + 0.40 × 0.60² = 0.24&lt;/code&gt;. For B, it is &lt;code&gt;0.60 × (0.95 − 1)² + 0.40 × 0.95² = 0.3625&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A's probabilities match the observed frequency in this constructed sample. B is overconfident — same direction accuracy, much worse calibration. The winner-only table cannot show that difference.&lt;/p&gt;

&lt;p&gt;One more caveat before you run off with this example: neither model distinguishes easy fixtures from hard ones; every probability is constant. A constant forecast can match the overall frequency while being useless at ranking individual matches.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use the right version of the score
&lt;/h2&gt;

&lt;p&gt;Football match direction normally has three classes — home win, draw, away win. A full probability forecast must assign each a non-negative probability, with the three summing to one.&lt;/p&gt;

&lt;p&gt;One common multiclass Brier convention averages the &lt;strong&gt;sum&lt;/strong&gt; of the squared errors across the three classes. Its range is zero to two. Other conventions rescale the score, so state your convention before comparing numbers.&lt;/p&gt;

&lt;p&gt;Log loss is another proper scoring rule. It punishes assigning very little probability to an outcome that actually occurs. If clipping is used to avoid numerical trouble at zero, disclose it — changing the clipping rule can change a reported score.&lt;/p&gt;

&lt;h2&gt;
  
  
  A reliability diagram needs counts
&lt;/h2&gt;

&lt;p&gt;To inspect calibration, group predictions into probability ranges and compare each group's mean predicted probability with its observed event rate. A group averaging 0.60 should see an observed rate near 0.60, subject to uncertainty.&lt;/p&gt;

&lt;p&gt;Publish the number of forecasts in each group. A point based on eight matches should not look as conclusive as one based on eight hundred.&lt;/p&gt;

&lt;p&gt;Scikit-learn's &lt;a href="https://scikit-learn.org/stable/modules/calibration.html" rel="noopener noreferrer"&gt;probability calibration guide&lt;/a&gt; explains reliability diagrams and adds an important caution: a lower Brier loss does not necessarily mean better calibration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare like with like
&lt;/h2&gt;

&lt;p&gt;A fair comparison uses the same held-out fixtures, the same publication horizon and the same result definition. Compare a twelve-hour forecast against a baseline that also had twelve hours of information — not against one updated after the team sheets came out.&lt;/p&gt;

&lt;p&gt;Look at the full sample and at sensible subgroups. A pooled reliability plot can hide overconfidence in one league and underconfidence in another.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calibration cannot be invented after the match
&lt;/h2&gt;

&lt;p&gt;Suppose an archived preview contains only "home win" plus a high-confidence badge. It's tempting to turn that badge into 80% after the event and compute a Brier score.&lt;/p&gt;

&lt;p&gt;That creates a brand-new forecast that was never published. Unless the numerical mapping was specified and recorded beforehand, it's not a valid probability-performance record.&lt;/p&gt;

&lt;p&gt;This applies to us too: our &lt;a href="https://xgmind.com/performance/" rel="noopener noreferrer"&gt;public results page&lt;/a&gt; should not be called a calibration dashboard merely because it displays direction and score outcomes. Probability metrics require the relevant pre-match probability vectors to exist and to have been preserved.&lt;/p&gt;

&lt;h2&gt;
  
  
  A short review routine
&lt;/h2&gt;

&lt;p&gt;Before accepting any model-comparison chart, check five things: the task, the held-out sample, the score convention, the forecast horizon, and the count behind each calibration point. Then read the errors rather than only the headline winner.&lt;/p&gt;

&lt;p&gt;And if a model looks unusually accurate, our &lt;a href="https://xgmind.com/blog/football-backtesting-data-leakage/" rel="noopener noreferrer"&gt;data-leakage guide&lt;/a&gt; is the next stop. Our &lt;a href="https://xgmind.com/methodology/" rel="noopener noreferrer"&gt;methodology&lt;/a&gt; documents how we keep that boundary intact.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;em&gt;Originally published at &lt;a href="https://xgmind.com/blog/football-forecast-calibration-brier-score/" rel="noopener noreferrer"&gt;xgmind.com/blog&lt;/a&gt; — the AI football match-intelligence lab with a fully public prediction ledger at &lt;a href="https://xgmind.com/performance/" rel="noopener noreferrer"&gt;xgmind.com/performance/&lt;/a&gt;.&lt;/em&gt;
&lt;/h2&gt;

</description>
      <category>machinelearning</category>
      <category>datascience</category>
      <category>statistics</category>
      <category>football</category>
    </item>
    <item>
      <title>The Problem With AI Football Predictions Isn't the AI</title>
      <dc:creator>sen yin</dc:creator>
      <pubDate>Fri, 02 Oct 2026 04:17:08 +0000</pubDate>
      <link>https://dev.to/sen_yin_caa8aa3b2b41be18f/the-problem-with-ai-football-predictions-isnt-the-ai-2egb</link>
      <guid>https://dev.to/sen_yin_caa8aa3b2b41be18f/the-problem-with-ai-football-predictions-isnt-the-ai-2egb</guid>
      <description>

&lt;p&gt;&lt;em&gt;By XG Mind AI (xgmind.com)&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A match preview lands in your feed with a confident 2–1 call, a fluent explanation of why the away captain's absence breaks the team's build-up play, and a link to the injury report it cites. Before you read a word of the tactical analysis, click the link. Does it name the same player, the same injury, the same competition, and the same fixture?&lt;/p&gt;

&lt;p&gt;That one click tells you more about the preview's worth than any debate about which language model wrote it. A beautiful explanation of the wrong absence is still wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fluency is not a data record
&lt;/h2&gt;

&lt;p&gt;Language models genuinely help with parts of this job: reading sources, writing analysis code, explaining what a model's outputs mean. But none of those abilities, on its own, gives a forecast predictive value. The question that actually matters is much narrower: &lt;strong&gt;where did the inputs come from, and what happened when this method was tested on matches it hadn't seen?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An ungrounded model response can lean on outdated squad news or stitch details together from different fixtures. Adding web search lowers some of that risk — but a citation is not verification. A report can link to a real news article while attributing that article's injury news to the wrong club.&lt;/p&gt;

&lt;p&gt;A careful reader should be able to pull a preview apart into three layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A reported fact&lt;/strong&gt;, like a club officially confirming a player's suspension.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A calculation&lt;/strong&gt;, like a scoring average over five named league fixtures.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An interpretation&lt;/strong&gt;, like the suggestion that the suspension might weaken defensive transitions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When those layers blur together, speculation starts wearing the uniform of a measured variable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Search is a research assistant, not a witness
&lt;/h2&gt;

&lt;p&gt;Search is genuinely useful here: it finds official team announcements, statistical tables, even downloadable datasets. Claiming it can never surface xG, or that text can never become a usable feature, would go too far.&lt;/p&gt;

&lt;p&gt;The hard part is consistency. Every search result needs the same provenance checks as any other source — identity, date, definition, coverage. Two sites can both publish a "last five games" table while counting different competitions. A page updated this morning may already include last night's fixture in its season totals, which breaks an analysis that claims to represent the previous afternoon.&lt;/p&gt;

&lt;p&gt;A few translations that help:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"Out for the weekend"&lt;/strong&gt; → Which weekend, which fixture, and was the announcement available before the preview was published?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Averaging 2.1 expected goals"&lt;/strong&gt; → Which provider, which season, which competition, what sample — and are penalties in the number?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Several outlets report"&lt;/strong&gt; → Independent outlets, or several copies of one original story?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"No injury news found"&lt;/strong&gt; → The squad is fit, or the search just didn't find anything?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one is the easiest to misread. An empty search result is evidence about the search, not proof of a fully fit squad. Our &lt;a href="https://xgmind.com/blog/missing-football-data-is-not-zero/" rel="noopener noreferrer"&gt;guide to missing data&lt;/a&gt; argues that distinction belongs inside the report itself, not hidden in the plumbing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Precise arithmetic can camouflage a vague assumption
&lt;/h2&gt;

&lt;p&gt;Take a &lt;strong&gt;hypothetical&lt;/strong&gt; independent-Poisson setup. One analyst feeds in expected goal rates of 2.1 and 0.8 and gets a score distribution. Another feeds in 1.6 and 1.2 and gets a different one.&lt;/p&gt;

&lt;p&gt;Both calculations can be arithmetically flawless. They disagree because the &lt;em&gt;assumptions&lt;/em&gt; disagree, and no number of decimal places in the output settles that disagreement.&lt;/p&gt;

&lt;p&gt;So for each rate estimate, ask what informed it: historical goals, a rolling xG sample, opponent adjustments — or an unsupported guess? Ask whether the method was fixed before it was tested. A spreadsheet, a Python script, or a language-model tool call should make this chain easier to inspect, not substitute for it.&lt;/p&gt;

&lt;p&gt;There's a second trap hiding here. Match xG is computed from shots that actually happened. A forecast of next Saturday's goal rate is a pre-match estimate. They're related ideas, but not interchangeable observations. Using Saturday's realised xG to "predict" Saturday's result leaks the match into its own forecast — a textbook case of what our &lt;a href="https://xgmind.com/blog/football-backtesting-data-leakage/" rel="noopener noreferrer"&gt;data-leakage guide&lt;/a&gt; warns against.&lt;/p&gt;

&lt;h2&gt;
  
  
  A reasonable division of labour
&lt;/h2&gt;

&lt;p&gt;Here's an architecture that works: a repeatable data process manages fixture identifiers, timestamps, metric definitions and missing values. A statistical procedure turns specified inputs into reproducible estimates. A language model helps inspect sources and communicate what those estimates mean.&lt;/p&gt;

&lt;p&gt;That's a useful architecture — not a certification stamp. A structured database can still hold stale or mis-mapped records. A statistical procedure can be badly specified. A language model can construct a persuasive causal story without any evidence that the proposed mechanism actually mattered.&lt;/p&gt;

&lt;p&gt;And not every serious project needs enterprise infrastructure. A carefully documented single-league study can be more reliable than a giant pipeline with weak checks. The standard is traceability and testing, not system size.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test comes after the explanation
&lt;/h2&gt;

&lt;p&gt;An attractive preview is not the evaluation. Preserve what was available before kickoff, record the forecast, and compare it with the result under rules chosen in advance. Separate match-direction accuracy from exact-score accuracy. And if probabilities are published, examine their calibration alongside their rankings.&lt;/p&gt;

&lt;p&gt;There is no universal accuracy ceiling that settles this — accuracy depends on the fixture mix and the task. A slate of heavy favourites is a different problem from a full league round full of coin-flip matches. Our piece on &lt;a href="https://xgmind.com/blog/the-mathematical-boundary-of-football-prediction/" rel="noopener noreferrer"&gt;the mathematical boundary of football prediction&lt;/a&gt; works through exactly that distinction.&lt;/p&gt;

&lt;p&gt;For us at XG Mind, the &lt;a href="https://xgmind.com/methodology/" rel="noopener noreferrer"&gt;methodology&lt;/a&gt; page and the &lt;a href="https://xgmind.com/performance/" rel="noopener noreferrer"&gt;public results ledger&lt;/a&gt; are starting points for scrutiny — not substitutes for it. Any claimed mechanism should be checked against what the platform actually publishes.&lt;/p&gt;

&lt;p&gt;So the best first question stays a plain one: can someone follow this number back to its source?&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;em&gt;Originally published at &lt;a href="https://xgmind.com/blog/why-most-ai-football-predictions-are-illusions/" rel="noopener noreferrer"&gt;xgmind.com/blog&lt;/a&gt; — the AI football match-intelligence lab with a fully public prediction ledger at &lt;a href="https://xgmind.com/performance/" rel="noopener noreferrer"&gt;xgmind.com/performance/&lt;/a&gt;.&lt;/em&gt;
&lt;/h2&gt;

</description>
      <category>machinelearning</category>
      <category>datascience</category>
      <category>football</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
