Every betting strategy sounds brilliant until you test it.
"Always back the home favourite." "Bet the draw when the match looks tight." "Fade the longshots." You can hear people argue about rules like these in any group chat. Almost nobody checks them against thousands of past matches.
In this tutorial we'll do exactly that. We'll build a complete backtesting pipeline in Python that:
Pulls multi-season historical results and closing odds from an API
Converts odds into implied probabilities and strips out the bookmaker's margin
Tests several strategies and reports ROI, hit rate and drawdown
Checks whether the results are statistically meaningful (or just luck)
Splits data by time, so you don't fool yourself with overfitting
I'll be honest up front: most simple strategies lose money once you test them properly. That isn't a failure of the tutorial. It's the main lesson. A backtest that tells you "don't bet this" has saved you real money.
Why Backtesting Matters
A backtest replays a rule against history and asks, "What would have happened if I had followed this rule mechanically?"
It's valuable because human memory is terrible at statistics. We remember the three times the underdog won and forget the forty times it didn't. A backtest forgets nothing.
It also teaches you the core ideas behind pricing and modelling:
Odds are prices with a margin built in. You pay that margin on every bet.
Returns are noisy. A strategy can look profitable over 100 bets and be pure luck.
The data you test on must not leak into the rule you design. This is called look-ahead bias, and it ruins more backtests than any bug.
Teams doing this professionally (trading desks, analytics teams, researchers) rely on exactly this workflow, just at larger scale.
What You'll Need
Python 3.10+
A free API key from Orbistats
Basic familiarity with pandas
About 45 minutes
Choosing a data source
Backtests are only as good as the history underneath them. The usual problems are shallow archives (two or three seasons), inconsistent team names, and a different format for history than for live data, so your model breaks the moment you deploy it.
I'm using the Orbistats Historical Sports Data API because it addresses those problems directly. It provides multi-season archives (football goes back to 2000+, most other sports to 2005 or 2010), including results, statistics, lineups and closing odds, all in the same schema as the live feed. That last point matters. A model trained on 2018 data reads identically to a match happening today, with no mapping layer in between.
Coverage differs by sport, so here's the depth the product page lists for all 13 sports:
Sport History from
Football 2000+
Basketball 2005+
American Football 2005+
Cricket 2005+
Tennis 2005+
Baseball 2010+
Combat Sports 2010+
Volleyball 2010+
Handball 2010+
Ice Hockey 2010+
Golf 2010+
Horse Racing 2010+
Esports 2015+
Treat that table as a ceiling, not a promise. Depth varies by league, and the coverage matrix in the API reference shows what's live per sport today. Check it before you pick a sport to test.
For this walkthrough I'll use English Premier League football, since it has the deepest archive and the cleanest three-way (home/draw/away) market.
Step 1: Get Your API Key and Explore
Create an account on the sign-up page and copy your key.
Read the documentation and the quickstart to see the auth and response conventions.
Try a few calls in the sandbox before writing code, so you can see the real JSON.
Authentication is a bearer token:
http
Authorization: Bearer YOUR_API_KEY
Base URL: https://api.orbistats.com/v1/
Every response uses a consistent envelope:
json
{
"data": [],
"meta": { "pagination": {}, "generated_at": "2026-08-28T10:00:00Z" },
"errors": []
}
A free key lets you explore recent seasons. Deeper archives, bulk export and higher limits sit on paid plans, so check the pricing page for what your tier includes.
Step 2: Project Setup
bash
mkdir odds-backtest && cd odds-backtest
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install requests pandas numpy matplotlib python-dotenv
Create .env:
env
ORBISTATS_API_KEY=your_key_here
And add it to .gitignore right now, before you forget:
text
.env
data_cache/
Step 3: Fetch Historical Odds and Results
Create backtest.py. First the configuration and the API client:
python
import os
import json
import time
import logging
from pathlib import Path
import numpy as np
import pandas as pd
import requests
import matplotlib.pyplot as plt
from dotenv import load_dotenv
load_dotenv()
API_KEY = os.environ["ORBISTATS_API_KEY"]
BASE_URL = "https://api.orbistats.com/v1"
SPORT = "football"
LEAGUE = "premier-league"
SEASON_FROM = "2019-2020"
SEASON_TO = "2024-2025"
The docs show two historical routes. Confirm which one your key uses
in the sandbox, then set it here:
"history/matches" (API reference)
"historical/results" (Historical Data page; season= instead of a range)
HISTORY_PATH = "history/matches"
CACHE = Path("data_cache") / f"{SPORT}{LEAGUE}{SEASON_FROM}_{SEASON_TO}.json"
CACHE.parent.mkdir(exist_ok=True)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger("backtest")
session = requests.Session()
session.headers.update({"Authorization": f"Bearer {API_KEY}"})
Now the fetcher. Historical queries return a lot of rows, so it needs pagination, and it needs to behave when it hits a rate limit:
python
def fetch_history(sport, league, season_from, season_to, per_page=100):
url = f"{BASE_URL}/{sport}/{HISTORY_PATH}"
params = {
"league": league,
"season_from": season_from,
"season_to": season_to,
"per_page": per_page,
}
rows, page = [], 1
while True:
resp = session.get(url, params=params, timeout=30)
if resp.status_code == 429:
wait = int(resp.headers.get("Retry-After", 30))
log.warning("Rate limited, sleeping %ss", wait)
time.sleep(wait)
continue
resp.raise_for_status()
body = resp.json()
batch = body.get("data") or []
rows.extend(batch)
log.info("Fetched %d rows (total %d)", len(batch), len(rows))
pagination = (body.get("meta") or {}).get("pagination") or {}
cursor = pagination.get("next_cursor")
if cursor: # cursor overrides page
params["cursor"] = cursor
elif batch and pagination.get("total") and len(rows) < pagination["total"]:
page += 1
params["page"] = page
else:
break
return rows
The reference recommends cursor pagination for large historical pulls and caps per_page at 100. Five Premier League seasons is about 1,900 matches, so that's roughly 19 requests. That's tiny compared with a 100-requests-per-minute limit, but it adds up fast if you loop over many leagues or sports.
Cache the raw response. You'll rerun your analysis dozens of times, and there's no reason to re-download history each time:
python
def load_raw():
if CACHE.exists():
log.info("Loading cached data from %s", CACHE)
return json.loads(CACHE.read_text())
raw = fetch_history(SPORT, LEAGUE, SEASON_FROM, SEASON_TO)
CACHE.write_text(json.dumps(raw))
return raw
Step 4: Turn Raw JSON into a Clean DataFrame
The API's sample match looks like this:
json
{
"fixture_id": 41207,
"home_team": "Arsenal",
"away_team": "Chelsea",
"final_score": { "home": 2, "away": 1 },
"closing_odds": { "home": 1.95, "draw": 3.60, "away": 3.80 },
"kickoff": "2024-10-27T15:30:00Z"
}
I always isolate the "messy to tidy" step in one function, so if a field name changes, there's exactly one place to edit:
python
def flatten(m: dict) -> dict | None:
try:
odds = m["closing_odds"]
score = m["final_score"]
return {
"fixture_id": m["fixture_id"],
"kickoff": m["kickoff"],
"home_team": m["home_team"],
"away_team": m["away_team"],
"home_goals": score["home"],
"away_goals": score["away"],
"odds_home": odds.get("home"),
"odds_draw": odds.get("draw"),
"odds_away": odds.get("away"),
}
except (KeyError, TypeError):
return None # skip matches with no odds or no final score
def build_frame(raw: list[dict]) -> pd.DataFrame:
rows = [r for r in map(flatten, raw) if r]
df = pd.DataFrame(rows)
df["kickoff"] = pd.to_datetime(df["kickoff"], utc=True)
for col in ("odds_home", "odds_draw", "odds_away"):
df[col] = pd.to_numeric(df[col], errors="coerce")
df = (
df.dropna(subset=["odds_home", "odds_away"])
.drop_duplicates("fixture_id")
.sort_values("kickoff")
.reset_index(drop=True)
)
df["result"] = np.select(
[df.home_goals > df.away_goals, df.home_goals < df.away_goals],
["home", "away"],
default="draw",
)
return df
Always sanity check before you trust anything:
python
df = build_frame(load_raw())
print(len(df), "matches")
print(df["kickoff"].min(), "->", df["kickoff"].max())
print(df["result"].value_counts(normalize=True).round(3))
For a Premier League sample you should see roughly 380 matches per season, and home wins should outnumber away wins. If you see 40 matches or wildly odd percentages, something is wrong with the fetch, not the strategy.
Step 5: Understand Odds, Implied Probability and the Margin
This is the most important concept in the whole tutorial, so let's slow down.
Decimal odds convert to an implied probability:
text
implied probability = 1 / odds
Odds of 2.00 imply 50%. Odds of 4.00 imply 25%. (The glossary has plain-English definitions of these and related terms if you want a refresher.)
Now take the sample match: home 1.95, draw 3.60, away 3.80.
text
1/1.95 = 0.513
1/3.60 = 0.278
1/3.80 = 0.263
Total = 1.054
The three probabilities add up to 105.4%, not 100%. That extra 5.4% is the bookmaker's margin (the "overround" or "vig"). It's the fee you pay on every bet.
To get margin-free "fair" probabilities, divide each implied probability by the total:
text
fair home = 0.513 / 1.054 = 48.7%
Here's the consequence for backtesting. If you blindly bet every outcome at a book with a 5% margin, your expected ROI isn't zero. It's about:
text
1 / 1.05 - 1 = -4.8%
That's your baseline. A strategy has to beat roughly minus the margin just to look average. Keep this number in your head when you read the results below.
Step 6: Reshape into One Row per Selection
Here's a design decision that pays off later. Instead of one row per match with three odds columns, build a long table with one row per selection (match plus outcome). Then every strategy becomes a simple filter, and the same code works for sports with two outcomes (tennis, basketball) or three.
python
OUTCOMES = ("home", "draw", "away")
def to_selections(df: pd.DataFrame) -> pd.DataFrame:
frames = []
for o in OUTCOMES:
col = f"odds_{o}"
part = df[["fixture_id", "kickoff", "home_team", "away_team", "result", col]]
part = part.rename(columns={col: "odds"}).assign(outcome=o)
frames.append(part)
sel = pd.concat(frames).dropna(subset=["odds"])
sel["implied"] = 1 / sel["odds"]
sel["overround"] = sel.groupby("fixture_id")["implied"].transform("sum")
sel["fair_p"] = sel["implied"] / sel["overround"]
sel["won"] = sel["outcome"] == sel["result"]
# rank 1 = shortest price = the favourite
sel["rank"] = sel.groupby("fixture_id")["odds"].rank(method="first")
sel["profit"] = np.where(sel["won"], sel["odds"] - 1, -1.0) # 1-unit flat stake
return sel.sort_values(["kickoff", "fixture_id"]).reset_index(drop=True)
sel = to_selections(df)
print("Average margin: {:.2%}".format(sel.groupby("fixture_id")["overround"].first().mean() - 1))
That last line tells you the typical margin in your dataset. Use it to compute your baseline ROI from Step 5.
Step 7: Define Strategies as Simple Rules
Each strategy is just a function that returns a boolean mask over the selections table:
python
def is_underdog(s):
return s["rank"] == s.groupby("fixture_id")["rank"].transform("max")
STRATEGIES = {
"Always home": lambda s: s["outcome"] == "home",
"Always draw": lambda s: s["outcome"] == "draw",
"Always away": lambda s: s["outcome"] == "away",
"Favourite": lambda s: s["rank"] == 1,
"Underdog": is_underdog,
"Longshots (p<12%)": lambda s: s["fair_p"] < 0.12,
"Tight games: draw": lambda s: (s["outcome"] == "draw") & (s["fair_p"] > 0.27),
}
These aren't meant to be winners. They're a spread of the kinds of rules people actually believe in, so you can see how the margin treats each one.
Step 8: Run the Backtest and Measure Results
python
def run_strategy(sel: pd.DataFrame, mask) -> pd.DataFrame:
bets = sel[mask(sel)].copy().sort_values("kickoff")
bets["cum"] = bets["profit"].cumsum()
return bets
def summarize(bets: pd.DataFrame) -> dict:
n = len(bets)
if n == 0:
return {"bets": 0}
peak = np.maximum(bets["cum"].cummax(), 0)
return {
"bets": n,
"hit_rate_%": round(100 * bets["won"].mean(), 1),
"avg_odds": round(bets["odds"].mean(), 2),
"profit_units": round(bets["profit"].sum(), 1),
"roi_%": round(100 * bets["profit"].mean(), 2),
"max_drawdown": round((bets["cum"] - peak).min(), 1),
}
results = {name: run_strategy(sel, rule) for name, rule in STRATEGIES.items()}
report = pd.DataFrame({name: summarize(b) for name, b in results.items()}).T
print(report.sort_values("roi_%", ascending=False))
How to read the table
bets: sample size. Be suspicious of any strategy with fewer than a few hundred.
hit_rate_%: how often it wins. Meaningless without the average odds next to it.
avg_odds: a 25% hit rate at average odds of 4.5 is a very different story from 25% at 2.5.
roi_%: profit divided by total staked. The number everyone cares about.
max_drawdown: the worst peak-to-trough fall in units. This is the pain you'd have to sit through, and it's what actually makes people abandon strategies.
Expect most rows to land near your baseline (minus the margin) with some noise. That's the market being reasonably efficient. If a strategy shows a big positive ROI, don't celebrate yet. Move to the next step.
Step 9: Is It Skill or Luck? Bootstrap the ROI
A strategy that made +6% over 300 bets might simply have gotten lucky. A bootstrap estimates how much the ROI could wobble by resampling your bets with replacement:
python
def bootstrap_roi(bets: pd.DataFrame, n_boot: int = 2000, seed: int = 42):
rng = np.random.default_rng(seed)
p = bets["profit"].to_numpy()
means = rng.choice(p, size=(n_boot, len(p)), replace=True).mean(axis=1) * 100
low, high = np.percentile(means, [2.5, 97.5])
return round(low, 2), round(high, 2)
for name, bets in results.items():
if len(bets) >= 100:
lo, hi = bootstrap_roi(bets)
print(f"{name:<22} 95% CI for ROI: [{lo:>6}%, {hi:>6}%]")
The rule of thumb: if the 95% interval includes zero (or your baseline), you cannot claim the strategy has an edge. Nearly every simple strategy fails this test, and learning to see that is half of what makes a good analyst.
Step 10: Check Calibration (Where Edges Actually Hide)
Instead of testing random rules, ask the data a sharper question: are the market's probabilities well calibrated? If outcomes priced at 10% really win 10% of the time, there's no free lunch. If they win 12%, there might be.
python
def calibration(sel: pd.DataFrame) -> pd.DataFrame:
bins = [0, .05, .10, .15, .20, .30, .40, .50, .60, .80, 1.0]
s = sel.copy()
s["bin"] = pd.cut(s["fair_p"], bins=bins)
return (
s.groupby("bin", observed=True)
.agg(n=("won", "size"),
fair_p=("fair_p", "mean"),
actual=("won", "mean"),
roi_pct=("profit", lambda x: 100 * x.mean()))
.round(3)
)
print(calibration(sel))
Compare fair_p to actual in each row. A well-known pattern in betting markets is the favourite-longshot bias: very long shots are often overpriced and heavy favourites slightly underpriced. Whether it shows up in your sample, and whether it survives the margin, is exactly the kind of thing a backtest can tell you. It's also a much better starting point for a model than a gut feeling.
Step 11: Avoid the #1 Backtesting Mistake (Overfitting)
Here's how people fool themselves. They try 50 rules, pick the one with the best ROI, and report it. But if you test 50 rules on the same data, a few will look great by pure chance.
The fix is simple: tune on the past, judge on the future.
python
cut = sel["kickoff"].sort_values().iloc[int(len(sel) * 0.6)]
train, test = sel[sel["kickoff"] <= cut], sel[sel["kickoff"] > cut]
best_thr, best_roi = None, -999
for thr in [0.06, 0.08, 0.10, 0.12, 0.15, 0.18]:
roi = 100 * train.loc[train["fair_p"] < thr, "profit"].mean()
print(f"train fair_p < {thr:.2f}: ROI {roi:6.2f}%")
if roi > best_roi:
best_thr, best_roi = thr, roi
test_roi = 100 * test.loc[test["fair_p"] < best_thr, "profit"].mean()
print(f"\nChosen threshold: {best_thr} train ROI {best_roi:.2f}% test ROI {test_roi:.2f}%")
If the test ROI collapses compared with the train ROI, you found noise, not an edge. That gap is the honest measure of how much a strategy is worth.
For something stricter, try walk-forward testing: tune on seasons 1 to 3, test on season 4, then tune on seasons 1 to 4, test on season 5, and so on. It mimics how you'd actually use the rule in real life.
Step 12: Visualize the Equity Curves
Numbers are great, but a picture of the drawdowns makes the risk feel real:
python
def plot_curves(results: dict):
fig, ax = plt.subplots(figsize=(11, 6))
for name, bets in results.items():
if len(bets):
ax.plot(bets["kickoff"], bets["cum"], label=name, linewidth=1.4)
ax.axhline(0, color="grey", linewidth=0.8)
ax.set_title("Cumulative profit by strategy (1-unit flat stakes)")
ax.set_ylabel("Profit (units)")
ax.legend(fontsize=8)
fig.tight_layout()
fig.savefig("equity_curves.png", dpi=150)
plot_curves(results)
You'll almost certainly see lines that drift downward at roughly the rate the margin predicts, with some noisy detours. Strategies that look profitable for six months and then fall off a cliff are the classic shape of a fluke.
Step 13: Simulate a Real Bankroll
Flat one-unit stakes are clean for analysis, but real money compounds. Here's a quick bankroll simulation staking a fixed percentage:
python
def simulate_bankroll(bets: pd.DataFrame, start=1000.0, pct=0.01) -> pd.Series:
bankroll, path = start, []
for _, b in bets.iterrows():
stake = bankroll * pct
bankroll += stake * (b["odds"] - 1) if b["won"] else -stake
path.append(bankroll)
return pd.Series(path, index=bets["kickoff"].values)
curve = simulate_bankroll(results["Favourite"])
print(f"Start 1000 -> End {curve.iloc[-1]:.0f}")
A note on Kelly staking: the formula f = (p*odds - 1) / (odds - 1) is popular, but it requires your own estimate of the true probability p. Plugging in the market's probability just gives you zero or negative stakes. Kelly is only as good as your model, and most people overestimate their model. Fractional Kelly (a quarter or half) is far safer.
Step 14: Run the Whole Thing
Add the entry point at the bottom of backtest.py:
python
if name == "main":
df = build_frame(load_raw())
sel = to_selections(df)
results = {n: run_strategy(sel, r) for n, r in STRATEGIES.items()}
report = pd.DataFrame({n: summarize(b) for n, b in results.items()}).T
print(report.sort_values("roi_%", ascending=False).to_string())
print("\nCalibration:")
print(calibration(sel).to_string())
plot_curves(results)
bash
python backtest.py
Pulling Much More Data: Bulk Export
REST pagination is fine for a few thousand matches. If you want ten seasons across several leagues, the historical product page documents a bulk export that returns whole seasons as a compressed file, instead of paging through thousands of calls. The documented request body looks like this:
http
POST /v1/football/historical/export
{
"league": "premier-league",
"seasons": ["2020", "2021", "2022", "2023", "2024"],
"include": ["results", "statistics", "odds"],
"format": "json"
}
Check your plan's access and the exact response behaviour in the docs before relying on it, since bulk export is listed as an upgraded feature.
Taking It to Other Sports
Because the historical data uses one schema for every sport, our to_selections design carries over with almost no changes. A few notes:
Two-way markets like tennis have no draw, so odds_draw is simply missing and the code drops it. Remember that "favourite" strategies behave very differently when the favourite wins 70% of the time.
Basketball is high-scoring and usually has no draw, so margins and calibration look different from football.
Cricket has formats (Test, ODI, T20) with different draw probabilities. Test each format separately rather than mixing them.
Coverage varies. The coverage matrix currently shows historical data for most sports, but not every one. Verify before building around a specific sport.
For the thirteen sports Orbistats lists, the same pipeline works wherever history exists: football, basketball, American football, cricket, tennis, baseball, esports, combat sports, volleyball, handball, ice hockey, golf and horse racing. Start with the deepest archives (football, basketball) while you learn.
Closing Line Value: A Better Test Than ROI
ROI over a few hundred bets is mostly noise. Professionals often use a different yardstick: closing line value (CLV). The idea is that the closing price is the market's most informed estimate. If you consistently bet at prices better than the close, you're probably finding real information, even before the profits show up.
To measure CLV you need an opening or intermediate price and the close. The historical odds documentation mentions closing lines on all plans and full line-movement history on higher plans, which is what makes this kind of analysis possible. If you're building something for a trading team, that's the data to look at; the Trading Desks page describes the use case, and the Analytics & Data Science page covers research-style workloads.
Common Backtesting Pitfalls (Read This Twice)
Look-ahead bias. Never use information that wasn't available at bet time. Final scores, end-of-season standings and post-match statistics are all off-limits for pre-match decisions.
Closing-odds optimism. You usually can't bet at the exact closing price, and stake limits and price changes apply. Treat closing-odds backtests as an upper bound.
Ignoring the margin. Always compare results to the baseline from Step 5, not to zero.
Tiny samples. Under a few hundred bets, almost anything can look good.
Data snooping. Every extra rule you try inflates the chance of a false discovery. Keep a held-out test set and touch it once.
Survivorship and coverage gaps. Check that every season has about the right number of matches, and look for leagues or periods with missing odds.
Ignoring drawdowns. A strategy with great ROI and a 60-unit drawdown is one you will probably abandon at the worst moment.
Ideas to Extend This Project
Fit a simple model (Elo ratings, Poisson goals, logistic regression) and compare its probabilities against the market's
Test Asian handicap, totals and over/under markets, not just 1X2
Add team form and statistics as features using the statistics endpoints
Compare strategies by season to see if an edge decays over time
Link this to a live system: backtest a rule, then feed it with the Odds API (note it requires an upgraded plan) so the live and historical data share the same schema
Wrapping Up
In one script we built a real backtesting pipeline:
Fetched multi-season results and closing odds via a paginated, rate-limit-aware client
Cleaned the data into a tidy DataFrame
Converted odds into implied and margin-free probabilities
Tested several strategies as simple filters
Measured ROI, hit rate and drawdown
Stress-tested the findings with bootstrap confidence intervals
Validated with a time-based train/test split
If you remember one thing, make it this: the goal of a backtest isn't to find a winner. It's to find out cheaply whether something deserves your trust. Most ideas won't. The few that survive an honest test are the only ones worth taking further.
If you build something on top of this (a walk-forward tester, a calibration dashboard, a multi-sport comparison), share it in the comments. And if the sandbox returns a different JSON shape than the sample I used, paste it below and I'll help adapt the flatten() function.
Happy testing! 📊
Top comments (1)
"A backtest that tells you don't bet this has saved you real money" is a much better framing than most tutorials open with.
Two traps worth adding a paragraph on, since you already cover look-ahead:
What does the time split look like in your pipeline: one train/test cut, or walk-forward by season?