<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Elio Liberatore</title>
    <description>The latest articles on DEV Community by Elio Liberatore (@commodus67).</description>
    <link>https://dev.to/commodus67</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4119915%2Fa4787c9b-3be1-48a1-b556-b475f5158135.png</url>
      <title>DEV Community: Elio Liberatore</title>
      <link>https://dev.to/commodus67</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/commodus67"/>
    <language>en</language>
    <item>
      <title>Playoff Probability Calibration — LLMs vs. a Real Monte Carlo Model</title>
      <dc:creator>Elio Liberatore</dc:creator>
      <pubDate>Sat, 26 Sep 2026 19:33:53 +0000</pubDate>
      <link>https://dev.to/commodus67/playoff-probability-calibration-llms-vs-a-real-monte-carlo-model-3j63</link>
      <guid>https://dev.to/commodus67/playoff-probability-calibration-llms-vs-a-real-monte-carlo-model-3j63</guid>
      <description>&lt;p&gt;&lt;em&gt;Submission for the DEV Community x Kaggle Benchmarking Challenge (#kagglechallenge)&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The question
&lt;/h2&gt;

&lt;p&gt;Large language models are increasingly asked to reason about probability — "what's the chance this team makes the playoffs?" is a perfectly natural thing to ask one. But do they actually &lt;em&gt;reason&lt;/em&gt; about it, or do they just echo whatever number is floating around in the sports media they were trained on?&lt;/p&gt;

&lt;p&gt;I already run a small sports-data business built on real Monte Carlo simulations (10,000-20,000 trials per team, cross-checked live against Kalshi prediction-market prices) for MLB and NFL playoff odds. That gave me a genuine, independently-verified ground truth to grade LLMs against — instead of grading them against each other, or against vibes.&lt;/p&gt;

&lt;p&gt;So I built a two-task Kaggle Benchmark, &lt;a href="https://www.kaggle.com/benchmarks/elioliberatore/playoff-probability-calibration-llms-vs-a-real-mo/leaderboard" rel="noopener noreferrer"&gt;&lt;strong&gt;"Playoff Probability Calibration: LLM vs Model"&lt;/strong&gt;&lt;/a&gt;, comparing 7 frontier and mid-tier models against my own simulation engine across 18 real MLB and NFL cases (5 MLB + 13 NFL).&lt;/p&gt;

&lt;h2&gt;
  
  
  The two tasks
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Task A — Probability calibration (pure reasoning).&lt;/strong&gt; The model gets the same plain-English context a bettor would have — team, record, games remaining, season point/run differential, a short narrative — and has to respond with a single number: its estimated playoff probability. No market price, no hints. Scored against my Monte Carlo model's own probability for that same team.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Task B — Monte Carlo code generation (agentic).&lt;/strong&gt; No opinions allowed here. The model has to &lt;em&gt;write a Python program&lt;/em&gt; that simulates the team's remaining games and prints a probability — and we actually execute what it writes. This isolates coding/agentic ability from sports "intuition": a model can't talk its way to a good score, it has to produce working simulation code that lands near the truth.&lt;/p&gt;

&lt;p&gt;Both tasks share the same 18-case dataset (&lt;code&gt;benchmark_dataset.csv&lt;/code&gt;), and 5 of those MLB cases resolve for real before the benchmark's own deadline — including a live Texas Rangers vs. Houston Astros game.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models tested
&lt;/h2&gt;

&lt;p&gt;Claude Opus 4.8, Claude Haiku 4.5, GPT-5.5, GPT-5.4 mini, Gemini 3.8 Flash, Gemini 3.7 Flash, and Qwen 3 Next 80B Instruct.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;(Mean score across all 18 cases, higher is better, 0-100%. See the live Kaggle benchmark for the underlying per-case breakdown.)&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Task A — Calibration&lt;/th&gt;
&lt;th&gt;Task B — Code-gen&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 mini&lt;/td&gt;
&lt;td&gt;98.0%&lt;/td&gt;
&lt;td&gt;99.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;97.3%&lt;/td&gt;
&lt;td&gt;99.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;97.1%&lt;/td&gt;
&lt;td&gt;99.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;95.4%&lt;/td&gt;
&lt;td&gt;99.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3 Next 80B Instruct&lt;/td&gt;
&lt;td&gt;&lt;em&gt;see note below&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;see note below&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;&lt;em&gt;see note below&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;see note below&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.8&lt;/td&gt;
&lt;td&gt;&lt;em&gt;see note below&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;&lt;em&gt;see note below&lt;/em&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;A note on the missing three.&lt;/strong&gt; Claude Opus 4.8, GPT-5.5, and Qwen 3 Next 80B Instruct could not complete either task on Kaggle's benchmarking infrastructure — every call to them failed before it was even billed, with a &lt;code&gt;403 PermissionDeniedError&lt;/code&gt;: &lt;em&gt;"The max estimated cost of operation exceeds your available quota (based on max_output_tokens)."&lt;/em&gt; I confirmed this is not simple quota exhaustion (it reproduces with most of the daily allowance still free) and not concurrency contention (it reproduces running a single model completely alone). There's no exposed way — in the task code, the model-selection UI, or the notebook itself — to lower the &lt;code&gt;max_output_tokens&lt;/code&gt; reservation these three models apparently require. I'm treating this as a hard platform-side limitation rather than a finding about the models themselves, and reporting the results here with the four models that could actually run.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually stood out
&lt;/h2&gt;

&lt;p&gt;Every model that &lt;em&gt;could&lt;/em&gt; run scored strikingly well — all four sit above 95% on both tasks, and the smaller/cheaper models (GPT-5.4 mini, the Gemini Flash pair) held their own right alongside the larger Claude Haiku 4.5. The bigger separation isn't between models — it's between the two &lt;em&gt;tasks&lt;/em&gt;: Task B (write-and-execute code) scores even higher and tighter across the board than Task A (state an opinion), which suggests these models are more reliable translating "simulate this" into working code than they are at directly reasoning their way to a well-calibrated number. That's a more interesting result than "model X beats model Y" — it says something about &lt;em&gt;how&lt;/em&gt; to prompt a model for a probability estimate at all: ask it to write the simulation, don't ask it to guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Live-resolving case: Texas Rangers vs. Houston Astros
&lt;/h2&gt;

&lt;p&gt;One of the 18 cases in the dataset is a real, still-upcoming MLB matchup — Texas Rangers vs. Houston Astros, around September 28, 2026. I'll update this section once the game resolves, with how each model's estimate compared to both the simulation and the actual outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  Methodology notes
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Ground truth: two of my own Apify-published Monte Carlo actors (&lt;code&gt;commodus67/mlb-playoff-odds-monte-carlo&lt;/code&gt;, &lt;code&gt;commodus67/nfl-playoff-odds-api-monte-carlo-simulator&lt;/code&gt;), each independently checked against live Kalshi prediction-market prices before being used as the benchmark's target.&lt;/li&gt;
&lt;li&gt;Built entirely on Kaggle's Benchmarks SDK (&lt;code&gt;kbench&lt;/code&gt;) — two &lt;code&gt;@kbench.task&lt;/code&gt;-decorated aggregate functions, one per task, each running the full 18-row dataset against whichever model Kaggle passes in.&lt;/li&gt;
&lt;li&gt;Every model got the same context, the same instructions, and the same scoring function — no reasoning traces or chain-of-thought hints, no tools beyond Python code execution for Task B.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it yourself
&lt;/h2&gt;

&lt;p&gt;The benchmark is set up to evaluate any additional model Kaggle supports — just point "Evaluate More Models" at either task. I'd be curious whether a model with more headroom on Kaggle's token-cost reservation (once I sort that out) closes the gap on the top four, or whether ~97-99% is close to a ceiling for this kind of task.&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>machinelearning</category>
      <category>sportsanalytics</category>
      <category>llm</category>
    </item>
    <item>
      <title>Projecting a tennis match while it's still being played, without point-by-point data</title>
      <dc:creator>Elio Liberatore</dc:creator>
      <pubDate>Thu, 24 Sep 2026 15:41:49 +0000</pubDate>
      <link>https://dev.to/commodus67/projecting-a-tennis-match-while-its-still-being-played-without-point-by-point-data-22h</link>
      <guid>https://dev.to/commodus67/projecting-a-tennis-match-while-its-still-being-played-without-point-by-point-data-22h</guid>
      <description>&lt;h2&gt;
  
  
  The problem: ESPN gives you a score, not a probability
&lt;/h2&gt;

&lt;p&gt;Public tennis data is thin compared to what's available for team sports. &lt;a href="https://site.api.espn.com/apis/site/v2/sports/tennis/atp/scoreboard" rel="noopener noreferrer"&gt;ESPN's tennis API&lt;/a&gt; returns the score of every ATP/WTA match, live or finished — sets won, current game, current point when it bothers to update — but nothing that looks like win probability, and no point-by-point log you could replay.&lt;/p&gt;

&lt;p&gt;That's a problem if you want to answer the question a live match actually raises: &lt;em&gt;given the score right now, who wins?&lt;/em&gt; A pre-match number computed from rankings goes stale the moment the first point is played. You need something that reacts to the score without needing data ESPN simply doesn't expose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with a number you can actually get: ranking points
&lt;/h2&gt;

&lt;p&gt;Before a ball is hit, the only signal worth trusting is each player's current ATP or WTA ranking points — public, updated weekly, and already a decent proxy for recent form. A &lt;a href="https://en.wikipedia.org/wiki/Bradley%E2%80%93Terry_model" rel="noopener noreferrer"&gt;Bradley-Terry model&lt;/a&gt; turns the ratio of two players' points into a pre-match win probability:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;p = pointsA^k / (pointsA^k + pointsB^k)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;k&lt;/code&gt; controls how much a ranking gap matters; &lt;code&gt;k = 1&lt;/code&gt; (plain ratio) turned out to minimize error against live Kalshi prices when I backtested it against five liquid ATP markets. A second parameter blends that raw probability toward a coin flip, because ranking points alone say nothing about surface, current form or head-to-head — the kind of thing a market prices in and a Bradley-Terry model on points cannot see.&lt;/p&gt;

&lt;p&gt;That gets you a solid pre-match number. It does not get you a live one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning a match probability into a set probability, backwards
&lt;/h2&gt;

&lt;p&gt;Here's the trick: a Grand Slam player up two sets to love is not simply "60% to win the match" scaled up. What actually changes when a set is won is the &lt;em&gt;number of sets still needed&lt;/em&gt;, not some vague match-level confidence score. So the real quantity to simulate isn't "the match" — it's "the sets that are left."&lt;/p&gt;

&lt;p&gt;That means I need a per-set win probability, not a per-match one. And I only have the match one.&lt;/p&gt;

&lt;p&gt;The fix is to invert it. For a best-of-three match, the probability of winning the match given a per-set probability &lt;code&gt;q&lt;/code&gt; has a closed form:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;P(match | q) = q² + 2·q²·(1 − q)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(win in two sets, or win the decider after splitting the first two.) I already know &lt;code&gt;P(match)&lt;/code&gt; from the Bradley-Terry step — call it &lt;code&gt;p&lt;/code&gt;. So I solve for &lt;code&gt;q&lt;/code&gt; such that &lt;code&gt;P(match | q) = p&lt;/code&gt;, by bisection: pick a &lt;code&gt;q&lt;/code&gt; in &lt;code&gt;[0, 1]&lt;/code&gt;, compute &lt;code&gt;P(match | q)&lt;/code&gt;, and narrow the interval until it converges on &lt;code&gt;p&lt;/code&gt;. A dozen iterations gets you machine precision; there's no need for anything fancier.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;invertSetProbability&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;matchProb&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;matchFn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;lo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;hi&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;mid&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;lo&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;hi&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;matchFn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;mid&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;matchProb&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;lo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;mid&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="nx"&gt;hi&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;mid&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;lo&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;hi&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Best-of-five gets the same treatment with its own (longer) closed form. Now I have a per-set probability that, run back through the formula, reproduces the pre-match number exactly — which means it's the right one to carry into the live simulation.&lt;/p&gt;

&lt;h2&gt;
  
  
  From set probability to a live score
&lt;/h2&gt;

&lt;p&gt;Once a match is actually in progress, the question becomes simple: from the &lt;em&gt;current&lt;/em&gt; score — sets won by each player — how many more sets does each need, and what's the chance of winning that many more out of what's left, at probability &lt;code&gt;q&lt;/code&gt; per set? That's just a binomial tail, and I compute it with a quick Monte Carlo simulation (thousands of trials of "flip the coin at probability &lt;code&gt;q&lt;/code&gt; until someone reaches the target") rather than deriving a closed form for every possible remaining-sets combination — simpler to write, and cheap enough to run per match.&lt;/p&gt;

&lt;p&gt;The result: a player already up one set shows a higher live probability than their pre-match number, without ever touching point-by-point data ESPN doesn't provide. The signal is coarse — set-level, not point-level — but it's honest about what it does and doesn't know, and it's free.&lt;/p&gt;

&lt;h2&gt;
  
  
  The format detail that almost broke best-of-five
&lt;/h2&gt;

&lt;p&gt;One wrinkle: ESPN's own &lt;code&gt;periods&lt;/code&gt; field, which should say whether a match is best-of-three or best-of-five, is unreliable — it reported 5 even for an ATP 250 event that's best-of-three. Trusting it would have inverted the wrong formula for most matches. The fix was to ignore it and detect the format from the tournament name instead: best-of-five only for the four ATP Grand Slams, best-of-three everywhere else — including the WTA majors, which are best-of-three regardless of prestige. A field that lies less than 50% of the time is worse than no field at all, because it's confident about being wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  A production crash that unit tests didn't catch
&lt;/h2&gt;

&lt;p&gt;The model and the live projection both checked out against static fixtures — 29 passing tests. The first real end-to-end run against live ESPN and Kalshi data crashed 55 matches in, with a schema validation error on an &lt;code&gt;undefined&lt;/code&gt; player name.&lt;/p&gt;

&lt;p&gt;The cause: ESPN's tennis scoreboard mixes doubles draws into the same feed as singles, undifferentiated by any simple flag. In a doubles match, each side is a team/roster object, not an &lt;code&gt;athlete&lt;/code&gt; — so the code path that reads &lt;code&gt;athlete.displayName&lt;/code&gt; returned &lt;code&gt;undefined&lt;/code&gt; for every doubles competitor, and &lt;code&gt;JSON.stringify&lt;/code&gt; silently drops &lt;code&gt;undefined&lt;/code&gt; values, which then failed a &lt;code&gt;required&lt;/code&gt; field check downstream. Unlike ESPN's own explicit &lt;code&gt;"TBD"&lt;/code&gt; placeholder for a genuinely unresolved future-round singles match, this wasn't a "no data yet" case — it was a shape mismatch that looked like missing data until you traced it back.&lt;/p&gt;

&lt;p&gt;The fix was two-fold: skip doubles groups (checked by both the group label and the shape of the competitor side, since either can be wrong on its own), and make sure the name-reading function always returns a string, falling back to &lt;code&gt;"TBD"&lt;/code&gt; only for a genuinely absent name. Two regression tests were added reproducing the exact draw shape that crashed. Static fixtures had been drawn from a single real match capture; the actual scoreboard on a random Tuesday has draws unit tests never saw. Rankings and prediction markets are singles-only anyway, so doubles rows were never wanted in the output — but "not wanted" and "safe to silently produce broken JSON for" are different bugs, and only one of them showed up until a real run hit it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this adds up to
&lt;/h2&gt;

&lt;p&gt;Nothing here needs anything ESPN doesn't already give away for free: rankings, a live score, set counts. The engineering is in turning a static pre-match number into something that reacts to a live score without inventing data that isn't there — bisection to get from "match probability" to "set probability," a small Monte Carlo to get from "set probability" back to "match probability, live" — and in not trusting a field just because it exists.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The model behind this is the &lt;a href="https://apify.com/commodus67/tennis-match-winner-monte-carlo" rel="noopener noreferrer"&gt;Tennis Match Winner Monte Carlo&lt;/a&gt; Actor — live and upcoming ATP/WTA singles matches, priced against Kalshi's KXATPMATCH/KXWTAMATCH prediction markets.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>statistics</category>
      <category>sports</category>
      <category>javascript</category>
    </item>
    <item>
      <title>NFL playoff odds after Week 1 (2026): why my Monte Carlo model barely trusts a 31-10 win</title>
      <dc:creator>Elio Liberatore</dc:creator>
      <pubDate>Tue, 15 Sep 2026 23:10:28 +0000</pubDate>
      <link>https://dev.to/commodus67/nfl-playoff-odds-after-week-1-2026-why-my-monte-carlo-model-barely-trusts-a-31-10-win-25k8</link>
      <guid>https://dev.to/commodus67/nfl-playoff-odds-after-week-1-2026-why-my-monte-carlo-model-barely-trusts-a-31-10-win-25k8</guid>
      <description>&lt;p&gt;Week 1 of the 2026 NFL season is done. Sixteen teams are 1-0, sixteen are 0-1, and every sports show is already telling you who is "for real."&lt;/p&gt;

&lt;p&gt;I build a Monte Carlo simulator that estimates NFL playoff odds for all 32 teams, so I wanted to know how much one week should move those numbers. The model's answer is "not much," and the reasons are more interesting than the numbers themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with one game of data
&lt;/h2&gt;

&lt;p&gt;A common way to estimate team strength is the Pythagorean expectation. You take points scored and points allowed and turn them into an expected win percentage. Across a full season it predicts future wins better than the win-loss record does.&lt;/p&gt;

&lt;p&gt;After one game it gives you nonsense. Here is what it says after Week 1, next to what my model actually uses:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Team&lt;/th&gt;
&lt;th&gt;Week 1&lt;/th&gt;
&lt;th&gt;Point diff&lt;/th&gt;
&lt;th&gt;Pythagorean (1 game)&lt;/th&gt;
&lt;th&gt;Model's strength estimate&lt;/th&gt;
&lt;th&gt;Playoff odds&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Jacksonville Jaguars&lt;/td&gt;
&lt;td&gt;W 34-10&lt;/td&gt;
&lt;td&gt;+24&lt;/td&gt;
&lt;td&gt;.948&lt;/td&gt;
&lt;td&gt;.679&lt;/td&gt;
&lt;td&gt;92.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kansas City Chiefs&lt;/td&gt;
&lt;td&gt;W 31-10&lt;/td&gt;
&lt;td&gt;+21&lt;/td&gt;
&lt;td&gt;.936&lt;/td&gt;
&lt;td&gt;.559&lt;/td&gt;
&lt;td&gt;69.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New York Jets&lt;/td&gt;
&lt;td&gt;W 23-10&lt;/td&gt;
&lt;td&gt;+13&lt;/td&gt;
&lt;td&gt;.878&lt;/td&gt;
&lt;td&gt;.411&lt;/td&gt;
&lt;td&gt;23.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Las Vegas Raiders&lt;/td&gt;
&lt;td&gt;W 27-13&lt;/td&gt;
&lt;td&gt;+14&lt;/td&gt;
&lt;td&gt;.850&lt;/td&gt;
&lt;td&gt;.399&lt;/td&gt;
&lt;td&gt;21.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New England Patriots&lt;/td&gt;
&lt;td&gt;L 10-13&lt;/td&gt;
&lt;td&gt;-3&lt;/td&gt;
&lt;td&gt;.349&lt;/td&gt;
&lt;td&gt;.597&lt;/td&gt;
&lt;td&gt;55.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Denver Broncos&lt;/td&gt;
&lt;td&gt;L 10-31&lt;/td&gt;
&lt;td&gt;-21&lt;/td&gt;
&lt;td&gt;.064&lt;/td&gt;
&lt;td&gt;.542&lt;/td&gt;
&lt;td&gt;39.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Los Angeles Rams&lt;/td&gt;
&lt;td&gt;L 7-27&lt;/td&gt;
&lt;td&gt;-20&lt;/td&gt;
&lt;td&gt;.039&lt;/td&gt;
&lt;td&gt;.544&lt;/td&gt;
&lt;td&gt;33.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Model run on September 15, 2026, after Week 1, with 20,000 simulated seasons. These numbers will change after Week 2.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two things stand out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Winning big barely helps if you weren't good before.&lt;/strong&gt; The Jets and Raiders both won by double digits and are still below 25%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Losing big barely hurts if you were.&lt;/strong&gt; The Patriots lost their opener and still have better playoff odds than the 1-0 Giants (39.9%). The Broncos lost by 21 and are still near 40%.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's by design. Here is how it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: shrink the current season hard
&lt;/h2&gt;

&lt;p&gt;The simulator estimates each team's true strength by blending three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The actual win percentage&lt;/li&gt;
&lt;li&gt;The Pythagorean expectation from point differential (weighted at 0.65 by default, because point differential is the better predictor)&lt;/li&gt;
&lt;li&gt;A baseline the estimate is pulled toward&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The pull is controlled by &lt;code&gt;regressionGames&lt;/code&gt;. Think of it as adding a number of imaginary games played at the baseline. The NFL default is 6.&lt;/p&gt;

&lt;p&gt;After Week 1, each team has one real game against six imaginary ones. So the current season is only about a seventh of the estimate. That's why a .936 Pythagorean number turns into something much closer to average.&lt;/p&gt;

&lt;p&gt;Six is much smaller than you'd use in baseball. A 17-game season means every game carries real information, and by midseason I want the real record to dominate. It just shouldn't dominate after one Sunday.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: don't regress everyone to .500
&lt;/h2&gt;

&lt;p&gt;If every team were pulled toward .500, all 32 would start the season with identical odds. That's useless, and it's obviously wrong.&lt;/p&gt;

&lt;p&gt;So the baseline comes from last season. The model takes last year's point differential, shrinks it toward .500 and uses that as each team's starting point. The &lt;code&gt;priorCarryover&lt;/code&gt; parameter sets how much survives. The default of 0.6 keeps 60% of last season's separation between teams.&lt;/p&gt;

&lt;p&gt;That's why Jacksonville's strength estimate is .679 even though most of the weight is still on the prior. It's also why a 21-point loss doesn't sink Denver.&lt;/p&gt;

&lt;p&gt;It also leads to this model's biggest weakness, and I'd rather say it upfront: &lt;strong&gt;it knows nothing about the offseason.&lt;/strong&gt; It doesn't know about a new quarterback, a new head coach or a key injury. The prior comes only from last season's scoreboard. Early in the season that's the main way the model can be wrong. The regression settings are there so real results take over quickly as the weeks pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: simulate the season 20,000 times
&lt;/h2&gt;

&lt;p&gt;With a strength estimate for each team, the simulator plays out every remaining game on the real schedule:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Strength uncertainty.&lt;/strong&gt; Each simulated season redraws every team's strength from a distribution (&lt;code&gt;strengthUncertainty&lt;/code&gt;, default 0.16). With a 17-game sample you can't be sure how good anyone is. If you set this to zero, the output becomes far too confident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Game results.&lt;/strong&gt; A logistic model decides each game from the strength gap. Home teams get +0.2 in log-odds, which works out to about a 55% win rate between evenly matched teams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ties.&lt;/strong&gt; Each game has a 0.4% chance of ending in a tie, roughly one per season across 272 games. As in the NFL, a tie counts as half a win.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Seeding.&lt;/strong&gt; Each simulated season is run through the NFL seeding rules: division winners, three wild cards per conference, and the No. 1 seed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After 20,000 seasons, the playoff odds are just how often each team made it. The same run returns division odds, wild-card odds, No. 1 seed odds and a projected final record. For example, the Chiefs project to about 9.7 wins right now and the Jaguars to about 11.8.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it yourself
&lt;/h2&gt;

&lt;p&gt;The simulator is published as an API on Apify: &lt;a href="https://apify.com/commodus67/nfl-playoff-odds-api-monte-carlo-simulator" rel="noopener noreferrer"&gt;NFL Playoff Odds API - Monte Carlo Simulator&lt;/a&gt;. Standings and the remaining schedule are loaded automatically. Every parameter above is an input, so you can test your own assumptions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"https://api.apify.com/v2/acts/commodus67~nfl-playoff-odds-api-monte-carlo-simulator/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "iterations": 20000,
    "regressionGames": 6,
    "pythagoreanWeight": 0.65,
    "priorCarryover": 0.6,
    "strengthUncertainty": 0.16,
    "homeFieldAdvantage": 0.2
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You get one row per team. Some settings worth playing with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;priorCarryover: 0&lt;/code&gt;&lt;/strong&gt; starts every team level and shows how little one week tells you on its own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;regressionGames&lt;/code&gt;&lt;/strong&gt; lower means you trust the 2026 results more, higher means you trust them less.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;season&lt;/code&gt;&lt;/strong&gt; runs a past season from where it stood, which is useful for backtesting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;archiveToNamedDataset&lt;/code&gt;&lt;/strong&gt; saves every run to a named dataset. Run it weekly and you get a history of how the odds moved, which you can't rebuild later.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can also pass your own market probabilities in &lt;code&gt;marketProbabilities&lt;/code&gt;, and the output will show where the model and the market disagree.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to watch next
&lt;/h2&gt;

&lt;p&gt;Week 4 is the next checkpoint worth looking at. By then a team has played four real games against six imaginary ones, and the current season starts to matter. If the model is still far from reality for a team that changed a lot in the offseason, that's a sign &lt;code&gt;priorCarryover&lt;/code&gt; should be lower in September.&lt;/p&gt;

&lt;p&gt;This is a statistical model, not a set of picks. Its job is to put a number on uncertainty, not to tell you who to back.&lt;/p&gt;

&lt;p&gt;If you like this kind of thing, I wrote about the same problem for hockey: &lt;a href="https://dev.to/commodus67/nhl-playoff-odds-for-2026-27-84-games-the-wild-card-rule-and-why-i-dont-trust-my-own-model-in-10hm"&gt;NHL playoff odds for 2026-27 and why I don't trust my own model in September&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>statistics</category>
      <category>sports</category>
      <category>api</category>
    </item>
    <item>
      <title>Shipping a statistical model to the browser: Dixon–Coles soccer predictions with no backend and no API key</title>
      <dc:creator>Elio Liberatore</dc:creator>
      <pubDate>Tue, 15 Sep 2026 15:50:06 +0000</pubDate>
      <link>https://dev.to/commodus67/shipping-a-statistical-model-to-the-browser-dixon-coles-soccer-predictions-with-no-backend-and-no-58p1</link>
      <guid>https://dev.to/commodus67/shipping-a-statistical-model-to-the-browser-dixon-coles-soccer-predictions-with-no-backend-and-no-58p1</guid>
      <description>&lt;p&gt;I wanted a public page where anyone could click a soccer match and see real probabilities behind it — 1X2, Over/Under, both teams to score, exact scorelines. The model already existed: a Dixon–Coles bivariate Poisson fitted on finished matches. The hard part was never the statistics. It was getting the thing onto a page that costs me nothing and exposes nothing.&lt;/p&gt;

&lt;p&gt;Here's what I ended up with, and why the obvious approaches don't work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three options that all fail
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Call the API from the browser.&lt;/strong&gt; Simplest to write, and it puts my API token in everyone's DevTools. Dead on arrival.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Put a proxy in front of it.&lt;/strong&gt; Now I'm running a server, and every visitor — including every bot — costs me a model run. A page that gets popular becomes a page that bills me for being popular. That's a bad shape for a free demo.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pre-render everything to static HTML.&lt;/strong&gt; Cheap and safe, but the whole point is interaction. I want you to flip the Over/Under line from 2.5 to 3.5 and see the number move. Pre-rendering every market for every line for every match is a combinatorial mess, and it's a lot of bytes to ship.&lt;/p&gt;

&lt;p&gt;All three fail for the same reason: they treat the model's &lt;em&gt;output&lt;/em&gt; as the thing to deliver.&lt;/p&gt;

&lt;h2&gt;
  
  
  The parameters are smaller than the answers
&lt;/h2&gt;

&lt;p&gt;A Dixon–Coles model, once fitted, is almost nothing. Per match, it's two expected goal rates — λ for the home side, λ for the away side. Per league, it's one low-score correlation parameter, ρ. That's it. Every market on the page is arithmetic over those three numbers.&lt;/p&gt;

&lt;p&gt;So don't ship the answers. Ship λh, λa, ρ, and let the browser do the arithmetic.&lt;/p&gt;

&lt;p&gt;The payload looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"generatedAt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-15T15:06:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Dixon-Coles bivariate Poisson, MLE-fitted attack/defence with exponential time decay (xi=0.0018), maxGoals=10"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"leagues"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"slug"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"eng.1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Premier League"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"rho"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;-0.115723&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"historyMatches"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1141&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"matches"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"k"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-18T19:00:00.000Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"h"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Brentford"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"a"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Chelsea"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"lh"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.670193&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"la"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.560041&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"q"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six keys per match, one of which (&lt;code&gt;q&lt;/code&gt;) is just a data-quality flag. A dozen leagues and a couple hundred upcoming fixtures fit in about 23 KB of JSON, inlined in the page. The whole document — markup, styles, logic, data — is around 54 KB. No fetch, no second request, no loading state.&lt;/p&gt;

&lt;h2&gt;
  
  
  The recompute is twenty lines
&lt;/h2&gt;

&lt;p&gt;Dixon–Coles is independent Poisson with a correction applied to the four low-scoring cells — 0–0, 0–1, 1–0, 1–1 — where real football deviates from independence. The correction is the τ function, and ρ is its single parameter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;grid&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;lh&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;la&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;rho&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;var&lt;/span&gt; &lt;span class="nx"&gt;g&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tau&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;MAXG&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;g&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;MAXG&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;pois&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;lh&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nf"&gt;pois&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;la&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="k"&gt;if      &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;tau&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;lh&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;la&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;rho&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;tau&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;lh&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;rho&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;tau&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;la&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;rho&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;tau&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;rho&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;else&lt;/span&gt;                         &lt;span class="nx"&gt;tau&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="nx"&gt;g&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;tau&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="nx"&gt;s&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;g&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;MAXG&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;MAXG&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;g&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;/=&lt;/span&gt; &lt;span class="nx"&gt;s&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;g&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An 11×11 grid of scorelines, renormalised so it sums to 1. Every market is then a sum over cells of that one grid:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;markets&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;var&lt;/span&gt; &lt;span class="nx"&gt;p1&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;px&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;p2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ov&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;bt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;MAXG&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;MAXG&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;v&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;g&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
      &lt;span class="k"&gt;if      &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="nx"&gt;p1&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;px&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;else&lt;/span&gt;              &lt;span class="nx"&gt;p2&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;ov&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;x&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;y&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;bt&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;p1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;px&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;px&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;p2&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;ov&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ov&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;un&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;ov&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;bt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;bt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;nb&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;bt&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the part I like most, and it's not a performance argument. Because every market comes out of the &lt;em&gt;same&lt;/em&gt; grid, the page is internally consistent by construction. The 1X2 probabilities, the Over/Under for every line, both-teams-to-score, and the top scorelines can't contradict each other, because they're five different sums over one object. If you've ever assembled a page like this from separate endpoints, you know how easy it is to publish a "draw" probability that disagrees with the sum of the 0–0, 1–1 and 2–2 cells sitting right next to it.&lt;/p&gt;

&lt;p&gt;Flipping the Over/Under line from 2.5 to 4.5 doesn't fetch anything. It changes one comparison in a loop over 121 cells.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check your reimplementation.&lt;/strong&gt; I ran the browser's grid against the scoreline grid the model itself returns, match by match. Largest disagreement: about 2e-7 — floating-point noise. Do this before you trust a client-side reimplementation of anything; "looks about right" is how you ship a subtly wrong model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping it fresh without a server
&lt;/h2&gt;

&lt;p&gt;A build script calls the model once per league, writes the parameters into the page template, and commits the result. GitHub Actions runs it on a cron; Pages serves it. Hosting cost is zero, per-visitor cost is zero, and the only recurring cost is one model run per league per day.&lt;/p&gt;

&lt;p&gt;Three things I got wrong first, so you don't have to:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The template lives inside the build script.&lt;/strong&gt; I edited the published HTML directly to fix something, felt good about it, and the next morning's run overwrote my fix. If a generator owns a file, the file is not the source. Obvious in hindsight; still cost me an afternoon.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fail loudly, publish conservatively.&lt;/strong&gt; If one league returns no fixtures, the script skips it. If &lt;em&gt;every&lt;/em&gt; league comes back empty, it exits non-zero and publishes nothing — so a bad upstream day leaves yesterday's good numbers up instead of replacing them with a blank page. The failure mode you want is stale, not empty.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scheduled doesn't mean punctual.&lt;/strong&gt; My cron says 11:30 UTC. Actual runs have landed three and a half and five and a half hours late. GitHub's scheduled workflows queue on shared capacity and there is no guarantee attached to that timestamp. If your copy says "updated every morning", your copy is wrong. Put the generation timestamp on the page and let it speak for itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Say when the model is thin
&lt;/h2&gt;

&lt;p&gt;Early in a European season, a promoted side has three or four matches of top-flight history. The fit still returns a number, and the number is confidently silly.&lt;/p&gt;

&lt;p&gt;So the payload carries a per-match quality flag, and the page renders a LOW DATA badge plus a plain-language card explaining that one of these teams has very little history behind its rate. It costs one flag per match in the payload and it's the difference between a page that reports numbers and a page that reports numbers honestly.&lt;/p&gt;

&lt;p&gt;Every model has a region where it shouldn't be trusted. Most interfaces hide it. Showing it costs almost nothing and is the single change most likely to make a technical reader believe the rest of your output.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it adds up to
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Model parameters, not model answers, on the wire&lt;/li&gt;
&lt;li&gt;~23 KB of JSON, one document, no runtime fetches&lt;/li&gt;
&lt;li&gt;No credentials in the client, because the client never calls anything&lt;/li&gt;
&lt;li&gt;Internally consistent markets, because they're sums over one grid&lt;/li&gt;
&lt;li&gt;Static hosting; cost doesn't scale with traffic&lt;/li&gt;
&lt;li&gt;Honest about staleness and about thin data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The technique generalises past football. Any fitted model whose parameters are small compared to its output surface can be shipped this way — pricing curves, survival models, anything where a handful of coefficients regenerate a large interactive result set. Ask what the smallest object is that lets the client rebuild the answer, and send that instead.&lt;/p&gt;

&lt;p&gt;Page: &lt;a href="https://commodus67.github.io/soccer-predictions-demo" rel="noopener noreferrer"&gt;commodus67.github.io/soccer-predictions-demo&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Model behind it: &lt;a href="https://apify.com/commodus67/soccer-dixon-coles-match-predictor" rel="noopener noreferrer"&gt;soccer-dixon-coles-match-predictor&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Simulation and data, not tips. Nothing here is betting advice.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>webdev</category>
      <category>showdev</category>
      <category>datascience</category>
    </item>
    <item>
      <title>NBA play-in odds are not playoff odds: simulating the play-in tournament instead of guessing it</title>
      <dc:creator>Elio Liberatore</dc:creator>
      <pubDate>Thu, 10 Sep 2026 22:26:19 +0000</pubDate>
      <link>https://dev.to/commodus67/nba-play-in-odds-are-not-playoff-odds-simulating-the-play-in-tournament-instead-of-guessing-it-1kae</link>
      <guid>https://dev.to/commodus67/nba-play-in-odds-are-not-playoff-odds-simulating-the-play-in-tournament-instead-of-guessing-it-1kae</guid>
      <description>&lt;p&gt;In baseball, American football and hockey, "make the playoffs" is one line: finish above it and you're in. When I built an NBA version of my playoff-odds simulator, I assumed the same, and my first plan was to compare the model against the wrong market.&lt;/p&gt;

&lt;p&gt;Basketball has a band in the middle, and getting it right changes almost every number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three zones, not two
&lt;/h2&gt;

&lt;p&gt;In each 15-team conference:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Finish&lt;/th&gt;
&lt;th&gt;What it means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1st – 6th&lt;/td&gt;
&lt;td&gt;Straight into the playoffs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7th – 10th&lt;/td&gt;
&lt;td&gt;Into the &lt;strong&gt;play-in tournament&lt;/strong&gt;, where two of four survive&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;11th – 15th&lt;/td&gt;
&lt;td&gt;Season over&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The play-in works like this. Seed 7 plays seed 8; the winner is the 7th seed. Seed 9 plays seed 10; the loser goes home. The loser of 7 v 8 then hosts the winner of 9 v 10, and that game decides the 8th seed.&lt;/p&gt;

&lt;p&gt;Kalshi lists these as separate markets — one for playoff qualification across all 30 teams, and one play-in market per conference — and its contract rules settle the question in one sentence: "Qualifying for the play-in tournament doesn't constitute playoff qualification."&lt;/p&gt;

&lt;p&gt;So the play-in market is not a weaker version of the playoff market. It's a band, and the two behave almost like opposites. A title contender is a near-certainty for the playoffs and a near-zero for the play-in. A 44-win team can be close to a coin flip between the two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Simulate the games, don't approximate them
&lt;/h2&gt;

&lt;p&gt;The tempting shortcut is to say "finish 7th or 8th, you're probably in; 9th or 10th, probably not". That breaks the moment you need two numbers that agree with each other: the chance of reaching the playoffs &lt;em&gt;and&lt;/em&gt; the chance of landing in the play-in.&lt;/p&gt;

&lt;p&gt;So in every simulated season the model:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Plays out the remaining regular-season schedule.&lt;/li&gt;
&lt;li&gt;Seeds each conference.&lt;/li&gt;
&lt;li&gt;Plays the three play-in games.&lt;/li&gt;
&lt;li&gt;Runs the full bracket: four rounds of best-of-seven series with the 2-2-1-1-1 home-court pattern.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That produces, for every team and from the same simulated seasons: top-six probability, play-in probability, playoff probability, conference finals, conference title and championship. Sanity checks come for free — the championship column sums to 1 across the league and conference finals to 4.&lt;/p&gt;

&lt;h2&gt;
  
  
  Basketball-specific modelling choices
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Pythagorean exponent.&lt;/strong&gt; Points scored and allowed predict future results better than record alone, but the exponent depends on the sport. Baseball uses about 1.83, hockey 2.0; for basketball I use 13.91. Plugging a baseball exponent into NBA scoring wildly compresses the gap between good and bad teams.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Uncertainty about strength itself.&lt;/strong&gt; My first version drew game outcomes from fixed team ratings. In September it confidently told me some teams made the playoffs 100% of the time and others 0%. That's not how the NBA works. Now each simulated season redraws every team's rating once (a 0.22 spread by default), so the output is a distribution rather than one confident guess. That alone cut the average gap against the market from 15.2 to 12.9 points and removed every 0% and 100%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The two missing games.&lt;/strong&gt; The 2026-27 regular season is 82 games, but ESPN's schedule currently lists 80 per team: the other two depend on the NBA Cup and are scheduled later. If you simply simulate the schedule you can see, every team's projected wins sit on an 80-game scale. I simulate the missing two against an average opponent on a neutral floor.&lt;/p&gt;

&lt;h2&gt;
  
  
  A September snapshot
&lt;/h2&gt;

&lt;p&gt;From one run on 10 September 2026:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Team&lt;/th&gt;
&lt;th&gt;Projected wins&lt;/th&gt;
&lt;th&gt;Top six&lt;/th&gt;
&lt;th&gt;Play-in&lt;/th&gt;
&lt;th&gt;Playoffs&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Oklahoma City&lt;/td&gt;
&lt;td&gt;55.3&lt;/td&gt;
&lt;td&gt;97.0%&lt;/td&gt;
&lt;td&gt;2.8%&lt;/td&gt;
&lt;td&gt;99.1%&lt;/td&gt;
&lt;td&gt;23.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;San Antonio&lt;/td&gt;
&lt;td&gt;52.1&lt;/td&gt;
&lt;td&gt;91.6%&lt;/td&gt;
&lt;td&gt;7.8%&lt;/td&gt;
&lt;td&gt;97.0%&lt;/td&gt;
&lt;td&gt;11.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boston&lt;/td&gt;
&lt;td&gt;51.6&lt;/td&gt;
&lt;td&gt;88.2%&lt;/td&gt;
&lt;td&gt;11.0%&lt;/td&gt;
&lt;td&gt;95.9%&lt;/td&gt;
&lt;td&gt;11.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toronto&lt;/td&gt;
&lt;td&gt;44.8&lt;/td&gt;
&lt;td&gt;51.0%&lt;/td&gt;
&lt;td&gt;41.0%&lt;/td&gt;
&lt;td&gt;73.7%&lt;/td&gt;
&lt;td&gt;2.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Atlanta&lt;/td&gt;
&lt;td&gt;44.4&lt;/td&gt;
&lt;td&gt;48.2%&lt;/td&gt;
&lt;td&gt;42.4%&lt;/td&gt;
&lt;td&gt;71.7%&lt;/td&gt;
&lt;td&gt;1.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Miami&lt;/td&gt;
&lt;td&gt;44.3&lt;/td&gt;
&lt;td&gt;47.9%&lt;/td&gt;
&lt;td&gt;42.6%&lt;/td&gt;
&lt;td&gt;71.3%&lt;/td&gt;
&lt;td&gt;1.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Look at Toronto, Atlanta and Miami: roughly a coin flip to avoid the play-in, and more than 40% to end up in it. For those teams, playoff odds and play-in odds are both "live" — exactly why they need separate columns.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number I don't trust
&lt;/h2&gt;

&lt;p&gt;Against Kalshi's playoff contracts that day, most of the table lines up within a few points: Oklahoma City 99.1% in the model against 98% on the market, Miami 71.3% against 72.5%, Toronto 73.7% against 70.5%.&lt;/p&gt;

&lt;p&gt;Then there's Charlotte: 86.5% in the model, 36% on the market. A fifty-point gap.&lt;/p&gt;

&lt;p&gt;That is not an opportunity. Before opening night, the model knows only how last season ended, regressed towards average. It hasn't seen free agency, the draft, trades or injuries. The market has. When I first measured the model against Kalshi in September, the rank correlation was about 0.80 and the average gap about 13 points — and the largest gaps were teams whose whole case this year is the offseason.&lt;/p&gt;

&lt;p&gt;So the simulator refuses to call anything "value" until every team has played ten games. It still reports every gap, labelled WATCH. A model that disagrees with the market by fifty points in September is telling you what it can't see, not what the market got wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  One market-side detail worth copying
&lt;/h2&gt;

&lt;p&gt;Playoff qualification is sixteen independent yes/no contracts, so prices across the league add up to about 16, not 1. Never normalise that field to 1. And price both sides of each contract: if the model says 60% and YES trades at 75 cents, the signal is on the NO side, not "no signal".&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The simulator is published as an Apify Actor: &lt;a href="https://apify.com/commodus67/nba-playoff-odds-monte-carlo" rel="noopener noreferrer"&gt;NBA Playoff Odds API — Monte Carlo Simulator &amp;amp; Value Bets&lt;/a&gt;. Standings and schedules come from ESPN and prices from Kalshi's public API, with no keys needed. Ready-made examples include &lt;a href="https://apify.com/commodus67/nba-playoff-odds-monte-carlo/examples/nba-play-in-tournament-odds-by-conference" rel="noopener noreferrer"&gt;play-in tournament odds by conference&lt;/a&gt;, &lt;a href="https://apify.com/commodus67/nba-playoff-odds-monte-carlo/examples/nba-championship-odds-simulated-playoff-bracket" rel="noopener noreferrer"&gt;championship odds from a simulated bracket&lt;/a&gt;, &lt;a href="https://apify.com/commodus67/nba-playoff-odds-monte-carlo/examples/nba-projected-win-totals-for-all-30-teams" rel="noopener noreferrer"&gt;projected win totals&lt;/a&gt; and &lt;a href="https://apify.com/commodus67/nba-playoff-odds-monte-carlo/examples/track-nba-playoff-odds-all-season" rel="noopener noreferrer"&gt;tracking playoff odds all season&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The season starts on 20 October. Simulation and data, not tips.&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>statistics</category>
      <category>sports</category>
      <category>simulation</category>
    </item>
    <item>
      <title>NHL playoff odds for 2026-27: 84 games, the wild card rule, and why I don't trust my own model in September</title>
      <dc:creator>Elio Liberatore</dc:creator>
      <pubDate>Thu, 10 Sep 2026 22:26:00 +0000</pubDate>
      <link>https://dev.to/commodus67/nhl-playoff-odds-for-2026-27-84-games-the-wild-card-rule-and-why-i-dont-trust-my-own-model-in-10hm</link>
      <guid>https://dev.to/commodus67/nhl-playoff-odds-for-2026-27-84-games-the-wild-card-rule-and-why-i-dont-trust-my-own-model-in-10hm</guid>
      <description>&lt;p&gt;I build Monte Carlo simulators that turn standings and schedules into playoff probabilities. After baseball, American football and soccer, I assumed hockey would be a copy-paste job with new team names.&lt;/p&gt;

&lt;p&gt;It wasn't. Four things about the NHL break a generic season simulator, and one of them changed this summer. Here they are, followed by the part I find most interesting: what the model says a few weeks before opening night, and why most of its disagreements with the market should be ignored.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Standings run on points, not wins
&lt;/h2&gt;

&lt;p&gt;A win is two points. An overtime or shootout loss is still worth one. Roughly 23% of NHL games go past regulation, so ranking simulated teams by wins misprices every club that lives in one-goal games.&lt;/p&gt;

&lt;p&gt;In the simulator each game has three outcomes — regulation win, overtime or shootout win, and the mirror images — and points are awarded the way the league awards them.&lt;/p&gt;

&lt;p&gt;There is a trap on the strength side too. Points percentage averages about .557 across the league because of the loser point, while win/loss averages exactly .500 because every game has a winner. If you estimate team strength from points percentage, every team looks slightly better than average. I estimate it from wins over games played, and from goals for and against with a Pythagorean exponent of 2.0.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The playoff field is not "the top eight"
&lt;/h2&gt;

&lt;p&gt;Sixteen of 32 teams qualify, but not by conference rank. In each conference the top three of each division get in, and then the two best remaining teams take the wild cards regardless of division. A fourth-place team in a strong division and a third-place team in a weak one are not interchangeable.&lt;/p&gt;

&lt;p&gt;That rule has to run inside every simulated season. It also gives you a free sanity check: across all 32 teams, playoff probabilities must sum to exactly 16. If yours sum to 15.7, something is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Overtime is closer to a coin flip
&lt;/h2&gt;

&lt;p&gt;Three-on-three and the shootout are not sixty minutes of five-on-five hockey. A better team carries less of its edge into the extra period. I damp the strength gap by half in overtime (&lt;code&gt;overtimeDamping = 0.5&lt;/code&gt;). Setting it to 0 makes overtime a pure coin flip; 1 treats it like regulation.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The season is 84 games now
&lt;/h2&gt;

&lt;p&gt;The collective bargaining agreement signed in 2025 moved the NHL to 84 regular-season games from 2026-27. ESPN's event count per team shows 84 for this season and 82 for last. I never hardcoded the season length — the simulator counts the games on the schedule feed — so projected point totals landed on the right scale without a code change. If you maintain a model with &lt;code&gt;82&lt;/code&gt; somewhere in it, now is the time to look.&lt;/p&gt;

&lt;p&gt;One more data quirk: ESPN names a season by the year it ends, so 2026-27 is &lt;code&gt;season=2027&lt;/code&gt;, and before opening night the 2027 standings tree exists but has zero teams in it. You need to fall back to the previous season for the list of clubs and divisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the model says right now
&lt;/h2&gt;

&lt;p&gt;Here is part of a 20,000-season run from 10 September 2026, next to the live price of Kalshi's "make the playoffs" contract for each team:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Team&lt;/th&gt;
&lt;th&gt;Projected points&lt;/th&gt;
&lt;th&gt;Model playoff %&lt;/th&gt;
&lt;th&gt;Kalshi&lt;/th&gt;
&lt;th&gt;Gap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Colorado&lt;/td&gt;
&lt;td&gt;109.6&lt;/td&gt;
&lt;td&gt;96.5%&lt;/td&gt;
&lt;td&gt;91%&lt;/td&gt;
&lt;td&gt;+5.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Carolina&lt;/td&gt;
&lt;td&gt;103.5&lt;/td&gt;
&lt;td&gt;82.9%&lt;/td&gt;
&lt;td&gt;89.5%&lt;/td&gt;
&lt;td&gt;−6.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Buffalo&lt;/td&gt;
&lt;td&gt;101.6&lt;/td&gt;
&lt;td&gt;76.3%&lt;/td&gt;
&lt;td&gt;56%&lt;/td&gt;
&lt;td&gt;+20.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edmonton&lt;/td&gt;
&lt;td&gt;95.7&lt;/td&gt;
&lt;td&gt;68.3%&lt;/td&gt;
&lt;td&gt;85%&lt;/td&gt;
&lt;td&gt;−16.7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vegas&lt;/td&gt;
&lt;td&gt;95.2&lt;/td&gt;
&lt;td&gt;66.4%&lt;/td&gt;
&lt;td&gt;85.5%&lt;/td&gt;
&lt;td&gt;−19.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Boston&lt;/td&gt;
&lt;td&gt;97.0&lt;/td&gt;
&lt;td&gt;58.3%&lt;/td&gt;
&lt;td&gt;31%&lt;/td&gt;
&lt;td&gt;+27.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pittsburgh&lt;/td&gt;
&lt;td&gt;95.5&lt;/td&gt;
&lt;td&gt;52.5%&lt;/td&gt;
&lt;td&gt;29.5%&lt;/td&gt;
&lt;td&gt;+23.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;San Jose&lt;/td&gt;
&lt;td&gt;89.7&lt;/td&gt;
&lt;td&gt;42.3%&lt;/td&gt;
&lt;td&gt;67.5%&lt;/td&gt;
&lt;td&gt;−25.2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Boston 27 points above the market. San Jose 25 below. If you believed the model, those would be the trades of the year.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I don't believe it (yet)
&lt;/h2&gt;

&lt;p&gt;In September the model knows exactly one thing about each club: how last season ended, shrunk 40% towards average. It hasn't seen a trade, a signing, an injury, a goalie change or a new coach. The market has seen all of them.&lt;/p&gt;

&lt;p&gt;So the biggest gaps in September are not edges. &lt;strong&gt;They are the offseason.&lt;/strong&gt; Betting them is betting that the summer didn't happen.&lt;/p&gt;

&lt;p&gt;When I first compared the model with Kalshi across all 32 teams, the rank correlation was 0.65 and the average gap was about 14 points, and the largest disagreements were precisely the teams whose summers changed the most. That is what a season-carryover model should look like before the puck drops.&lt;/p&gt;

&lt;p&gt;Instead of hiding that, I made it a rule in the output. Until every team has played a minimum number of games (10 by default), no row is allowed to call itself value. Every gap is still reported in full, but it's labelled WATCH. After that, the current season's record takes over from last season's gradually, rather than overnight after a hot week.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fee changes which gaps matter
&lt;/h2&gt;

&lt;p&gt;Kalshi's taker fee is 0.07 × p × (1 − p) per contract. It peaks in the middle: 1.75 cents on a 50-cent contract, 0.63 cents on a 90-cent one. So a 1.5-point edge on a coin-flip contract is a losing position after fees, while the same gap on a heavy favourite isn't. Any comparison that ignores this will rank the wrong teams first.&lt;/p&gt;

&lt;p&gt;A second trap: the playoff market is sixteen independent yes/no contracts, so prices across the league add up to about 16, not 1. If you "de-vig" it the way you would a division-winner market, you divide every probability by sixteen and manufacture huge fake edges everywhere. Division winners, on the other hand, are exclusive — exactly one team wins — and there you do strip the overround.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching it move
&lt;/h2&gt;

&lt;p&gt;The part I'm looking forward to is not the September table. It is the path. Each run can append its 32 rows to a named dataset that keeps growing, so a daily schedule gives you, by spring, how each team's probability moved across the season next to what the market charged for it on the same day — something you can't reconstruct afterwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The simulator is published as an Apify Actor: &lt;a href="https://apify.com/commodus67/nhl-playoff-odds-monte-carlo" rel="noopener noreferrer"&gt;NHL Playoff Odds API — Monte Carlo Simulator &amp;amp; Value Bets&lt;/a&gt;. No API keys for the data; standings and schedules come from ESPN, prices from Kalshi's public API. There are ready-made examples for &lt;a href="https://apify.com/commodus67/nhl-playoff-odds-monte-carlo/examples/nhl-projected-points-standings-all-32-teams" rel="noopener noreferrer"&gt;projected points standings&lt;/a&gt;, &lt;a href="https://apify.com/commodus67/nhl-playoff-odds-monte-carlo/examples/nhl-wild-card-odds-all-32-teams" rel="noopener noreferrer"&gt;wild card odds&lt;/a&gt;, &lt;a href="https://apify.com/commodus67/nhl-playoff-odds-monte-carlo/examples/nhl-presidents-trophy-odds" rel="noopener noreferrer"&gt;Presidents' Trophy odds&lt;/a&gt; and &lt;a href="https://apify.com/commodus67/nhl-playoff-odds-monte-carlo/examples/track-nhl-playoff-odds-all-season" rel="noopener noreferrer"&gt;tracking playoff odds all season&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Simulation and data, not tips.&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>statistics</category>
      <category>sports</category>
      <category>simulation</category>
    </item>
    <item>
      <title>Correct score, BTTS and Over/Under probabilities with Dixon-Coles: what I learned building it for MLS and Liga MX</title>
      <dc:creator>Elio Liberatore</dc:creator>
      <pubDate>Thu, 10 Sep 2026 22:25:35 +0000</pubDate>
      <link>https://dev.to/commodus67/correct-score-btts-and-overunder-probabilities-with-dixon-coles-what-i-learned-building-it-for-39dc</link>
      <guid>https://dev.to/commodus67/correct-score-btts-and-overunder-probabilities-with-dixon-coles-what-i-learned-building-it-for-39dc</guid>
      <description>&lt;p&gt;Most football prediction pages give you three numbers — home, draw, away — and no way to check where they came from. I wanted the opposite: one model, fitted on real results, that produces the 1X2 probabilities, the Over/Under 2.5 line, Both Teams To Score and the exact-score grid, all from the same place, so they can't contradict each other.&lt;/p&gt;

&lt;p&gt;The model I ended up with is Dixon-Coles. This post is what it does, why it beats the simpler version most tutorials start with, and what came out when I ran it on eight leagues that don't get much attention from modellers: MLS, Liga MX, Liga de Expansión MX, the Brasileirão Série B, the USL Championship, Colombia's Primera A, Uruguay's Primera División and Norway's Eliteserien.&lt;/p&gt;

&lt;h2&gt;
  
  
  The starting point: independent Poisson
&lt;/h2&gt;

&lt;p&gt;The classic approach gives every team an attack rating and a defence rating, adds a home advantage, and turns them into expected goals for each side of a fixture — call them λ for the home team and μ for the away team. Goals are then treated as two independent Poisson variables. The probability of a 2-1 is just P(home scores 2) × P(away scores 1).&lt;/p&gt;

&lt;p&gt;It works surprisingly well. It also has one known blind spot: &lt;strong&gt;low scores&lt;/strong&gt;. Real matches finish 0-0 and 1-1 more often than two independent Poisson draws predict, and 1-0 / 0-1 slightly less often. Anything that depends on those four cells — the draw, Under 2.5, BTTS "No" — inherits the error.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Dixon-Coles correction
&lt;/h2&gt;

&lt;p&gt;In 1997 Mark Dixon and Stuart Coles proposed a small fix. Keep the Poisson grid, but multiply the four low-score cells by a factor that depends on one extra parameter, ρ (rho):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;th&gt;Adjustment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0-0&lt;/td&gt;
&lt;td&gt;1 − λμρ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0-1&lt;/td&gt;
&lt;td&gt;1 + λρ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1-0&lt;/td&gt;
&lt;td&gt;1 + μρ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1-1&lt;/td&gt;
&lt;td&gt;1 − ρ&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;anything else&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With a negative ρ, 0-0 and 1-1 go up and 1-0 / 0-1 go down — exactly the direction the data pulls.&lt;/p&gt;

&lt;p&gt;To see how much that matters, take a match with 1.35 expected goals for the home side and 1.15 for the away side, and an illustrative ρ of −0.13:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Independent Poisson&lt;/th&gt;
&lt;th&gt;Dixon-Coles&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Draw&lt;/td&gt;
&lt;td&gt;26.8%&lt;/td&gt;
&lt;td&gt;30.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0-0&lt;/td&gt;
&lt;td&gt;8.2%&lt;/td&gt;
&lt;td&gt;9.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1-1&lt;/td&gt;
&lt;td&gt;12.7%&lt;/td&gt;
&lt;td&gt;14.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home win&lt;/td&gt;
&lt;td&gt;41.3%&lt;/td&gt;
&lt;td&gt;39.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Away win&lt;/td&gt;
&lt;td&gt;31.8%&lt;/td&gt;
&lt;td&gt;30.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same expected goals, three and a half points more on the draw. If you compare model probabilities with prices, that is the difference between seeing an edge and not seeing one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things the paper adds that tutorials often skip
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Time decay.&lt;/strong&gt; A result from three seasons ago should not count as much as last weekend's. Dixon and Coles weight each match by exp(−ξ·t), where t is its age in days. I use ξ = 0.0018, which halves a result's weight after roughly a year, and fit on three seasons of history.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fit ρ from the league itself.&lt;/strong&gt; ρ is not a universal constant. After fitting attack, defence, home advantage and the baseline by weighted maximum likelihood, I fit ρ separately for each league. They come out different. On 6 September, Colombia's Primera A gave a home advantage of 0.349 (on the log scale) and ρ = −0.044; Norway's Eliteserien gave 0.276 and ρ = −0.015. Colombian home sides get a bigger boost, and the low-score correction matters less in Norway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the data comes from
&lt;/h2&gt;

&lt;p&gt;All eight leagues come from ESPN's public scoreboard endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://site.api.espn.com/apis/site/v2/sports/soccer/&amp;lt;league&amp;gt;/scoreboard?dates=YYYYMMDD-YYYYMMDD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things cost me time and might save you some:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ask for a date range, not a single day.&lt;/strong&gt; With &lt;code&gt;dates=&lt;/code&gt; set to one day, a smaller league often returns nothing simply because it didn't play that day, which looks exactly like missing data. A week-long range removes the ambiguity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Filter by state, not by status name.&lt;/strong&gt; Finished matches carry &lt;code&gt;status.type.state === "post"&lt;/code&gt; and &lt;code&gt;completed: true&lt;/code&gt;; scheduled ones are &lt;code&gt;"pre"&lt;/code&gt;. That is more robust than matching &lt;code&gt;STATUS_FULL_TIME&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Slugs are not always what you'd guess. The USL Championship is &lt;code&gt;usa.usl.1&lt;/code&gt;, not &lt;code&gt;usa.2&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;History depth is not the bottleneck. The Premier League scoreboard answers back to at least 2002-03, and MLS and Liga MX back to at least 2004-05. Three seasons is plenty.&lt;/p&gt;

&lt;h2&gt;
  
  
  What one prediction looks like
&lt;/h2&gt;

&lt;p&gt;Here is a real row from a test run on 6 September — Atlanta United at home to Orlando City in MLS:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;λ home / λ away&lt;/td&gt;
&lt;td&gt;1.57 / 1.63&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ρ (MLS)&lt;/td&gt;
&lt;td&gt;−0.040&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Home / Draw / Away&lt;/td&gt;
&lt;td&gt;36.5% / 24.2% / 39.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Over 2.5 goals&lt;/td&gt;
&lt;td&gt;62.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Both teams to score&lt;/td&gt;
&lt;td&gt;64.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And from Liga MX the same day, Pumas UNAM against León: λ 1.94 against 0.94, so 60.1% / 23.0% / 16.9%. A strong home side, and the numbers say so.&lt;/p&gt;

&lt;p&gt;Every probability comes from one score grid (0-0 up to 10-10), normalised to 1, so 1X2 sums to 1, Over + Under sums to 1 and BTTS Yes + No sums to 1. I checked that on every row of a 28-match Série B run; the error was floating-point noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  The awkward case: promoted teams
&lt;/h2&gt;

&lt;p&gt;A team with no matches in the lookback window has no rating. The honest options are to guess or to say so. I start it at league-average strength and flag every fixture it plays with &lt;code&gt;dataQuality: "partial-new-team"&lt;/code&gt;, so whoever uses the numbers can decide how much to trust them until the team has a few games on the board.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't do
&lt;/h2&gt;

&lt;p&gt;It doesn't know about injuries, suspensions, rotation, weather or a manager who has just been sacked. It treats every match in the lookback window the same apart from its age. It is a baseline, not an oracle — which is exactly what makes it useful to compare against prices.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it without writing the model
&lt;/h2&gt;

&lt;p&gt;I packaged all of this as an Apify Actor: &lt;a href="https://apify.com/commodus67/soccer-dixon-coles-match-predictor" rel="noopener noreferrer"&gt;Soccer Match Predictions API — 1X2, Over/Under &amp;amp; BTTS Odds&lt;/a&gt;. You pick a league, it returns one row per upcoming fixture with everything above. There are ready-made examples, such as &lt;a href="https://apify.com/commodus67/soccer-dixon-coles-match-predictor/examples/mls-correct-score-probabilities" rel="noopener noreferrer"&gt;MLS correct score probabilities&lt;/a&gt; and &lt;a href="https://apify.com/commodus67/soccer-dixon-coles-match-predictor/examples/liga-mx-match-predictions-1x2-over-under-btts" rel="noopener noreferrer"&gt;Liga MX match predictions&lt;/a&gt;, and the rest are listed on the &lt;a href="https://apify.com/commodus67/soccer-dixon-coles-match-predictor/examples" rel="noopener noreferrer"&gt;examples page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you'd rather build it yourself, the table of adjustments above and the ESPN endpoint are all you need to get started. Either way: simulation and data, not tips.&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>statistics</category>
      <category>sports</category>
      <category>api</category>
    </item>
  </channel>
</rss>
