DEV Community

Cover image for Build a +EV Bet Finder in Python Using a Free Sports Odds API
orbistats
orbistats

Posted on

Build a +EV Bet Finder in Python Using a Free Sports Odds API

About 150 lines of Python, a mock mode that runs with no API key, and the math you need to understand every number it prints.

Disclosure: This is an educational engineering tutorial, not betting advice. A positive expected value on paper is not a guaranteed profit, and you can lose money. 18+ only, and check the laws where you live.

What We're Building

A command-line tool that:

Pulls odds for an event from several bookmakers.
Removes each bookmaker's margin (the "vig") to estimate the true probability.
Flags any price that beats that fair probability by more than a threshold (+EV, positive expected value).
Suggests a conservative stake with fractional Kelly.
Logs picks so you can measure closing line value (CLV), the honest scorecard for this kind of tool.

You will end up with this layout:

ev-finder/
├── fetch.py # API calls, caching, rate-limit handling, normalization
├── ev.py # de-vig, fair probabilities, EV, Kelly
├── main.py # CLI, alerts
├── tracker.py # SQLite pick log + CLV
├── mock_odds.json # sample data so it runs without a key
└── test_ev.py
The Idea in 60 Seconds

A bookmaker's odds contain a margin. If you remove it, the price implies a probability. If one bookmaker offers a better price than a reference probability justifies, the bet has positive expected value:

EV = fair_probability × decimal_odds − 1

A fair probability of 48.9% at odds of 2.20 gives 0.489 × 2.20 − 1 = +7.6%. The catch is the phrase "fair probability." Everything depends on how good your reference is. Sharp books (low margin, high limits) are the usual reference, and if none is available, an average across many books is a reasonable fallback. Background reading: expected value, vigorish, and the Kelly criterion.

Step 1: Pick a Data Source

You need decimal odds from several bookmakers on the same event, ideally with a free tier so you can prototype. Options include Orbistats, which lists 13 sports (Football, Basketball, American Football, Cricket, Tennis, Baseball, Esports, Combat Sports, Volleyball, Handball, Ice Hockey, Golf and Horse Racing) and describes normalized odds across bookmakers on its odds API page, and The Odds API, which focuses on odds. Larger commercial options exist too, such as SportsDataIO.

For Orbistats:

Create a key (free tier per their tutorials) and read the documentation.
Look at real response shapes in the public sandbox before writing any parsing code.
The base URL is https://api.orbistats.com/v1/ with a Bearer token, and the docs describe a sport → resource → endpoint pattern (for example /v1/football/fixtures). I assume /v1/{sport}/odds below. Confirm the exact path, parameters, and response fields in the docs and sandbox.

Check the live pricing page for free-tier limits, and read the changelog so an API change doesn't surprise you.

Step 2: Setup
bash
mkdir ev-finder && cd ev-finder
python -m venv .venv && source .venv/bin/activate
pip install requests python-dotenv pytest
echo "ORBISTATS_API_KEY=your_key_here" > .env
echo "DISCORD_WEBHOOK_URL=" >> .env # optional
echo ".env" >> .gitignore # never commit keys

Libraries: Requests, python-dotenv, pytest.

Step 3: The Math Module (ev.py)

This is the heart of the project, and it has no network code, so it's easy to test.

python

ev.py

from dataclasses import dataclass

def devig_power(odds_list):
"""Remove the margin with the power method.

Finds k such that sum((1/odds)**k) == 1. Compared with simple
proportional scaling, it moves more of the margin onto longshots,
which matches the favorite-longshot bias seen in many markets.
"""
inv = [1.0 / o for o in odds_list]
lo, hi = 0.5, 5.0
for _ in range(100):                      # fixed iterations, no infinite loop
    k = (lo + hi) / 2
    if sum(i ** k for i in inv) > 1:
        lo = k
    else:
        hi = k
return [i ** k for i in inv]
Enter fullscreen mode Exit fullscreen mode

def devig(prices):
"""{'home': 2.0, 'draw': 3.6, 'away': 3.9} -> fair probabilities dict."""
sels = list(prices)
return dict(zip(sels, devig_power([prices[s] for s in sels])))

def kelly_fraction(odds, p):
b = odds - 1.0
return max(0.0, (p * b - (1.0 - p)) / b)

def reference_probs(books, sharp_books, min_books=3):
"""Return (fair_probs, reference_name) for one event.

1) Use the first sharp book that quotes a complete market.
2) Otherwise average the de-vigged probabilities of all complete books.
"""
n = max(len(p) for p in books.values())
full = {b: p for b, p in books.items() if len(p) == n}
for sb in sharp_books:
    if sb in full:
        return devig(full[sb]), sb
if len(full) >= min_books:
    per_book = [devig(p) for p in full.values()]
    avg = {s: sum(d[s] for d in per_book) / len(per_book) for s in per_book[0]}
    total = sum(avg.values())
    return {s: v / total for s, v in avg.items()}, "consensus"
return None, None
Enter fullscreen mode Exit fullscreen mode

@dataclass
class Pick:
event: str
start: str
book: str
selection: str
odds: float
fair_prob: float
ev: float
kelly: float
ref: str

def find_value(events, sharp_books=(), min_ev=0.02, max_ev=0.15,
kelly_mult=0.25, min_books=3):
picks = []
for ev in events.values():
fair, ref = reference_probs(ev["books"], sharp_books, min_books)
if not fair:
continue
for book, prices in ev["books"].items():
if book == ref:
continue # don't compare a book to itself
for sel, odds in prices.items():
if sel not in fair:
continue
edge = fair[sel] * odds - 1.0
# Very large "edges" are usually stale prices or bad data. Skip them.
if min_ev <= edge <= max_ev:
picks.append(Pick(ev["label"], ev["start"], book, sel, odds,
fair[sel], edge,
kelly_fraction(odds, fair[sel]) * kelly_mult,
ref))
return sorted(picks, key=lambda p: p.ev, reverse=True)

Two design choices worth understanding:

max_ev cap: when a "+25% edge" appears, the likelier explanation is a stale line, a suspended market, or a mismatched selection, not free money. Treat it as a data-quality alert.
kelly_mult=0.25: full Kelly assumes your probabilities are exactly right. They never are, so quarter Kelly is a common way to reduce the damage from model error.
Step 4: Fetching and Normalizing Odds (fetch.py)

Real APIs use different field names, so the rest of the app only ever sees one internal shape:

python
{ event_id: {"label": "A vs B", "start": "...", "books": {"bookName": {"home": 2.1, "draw": 3.4, "away": 3.6}}} }
python

fetch.py

import json, os, time
import requests

BASE = "https://api.orbistats.com/v1"
_cache = {}

def _get(url, headers, ttl=30, params=None):
key = (url, tuple(sorted((params or {}).items())))
hit = _cache.get(key)
if hit and time.time() - hit[0] < ttl:
return hit[1] # serve from cache, save your quota
delay = 1
for _ in range(5):
r = requests.get(url, headers=headers, params=params, timeout=10)
if r.status_code == 429: # honor Retry-After when it's a number
wait = r.headers.get("Retry-After", "")
time.sleep(int(wait) if wait.isdigit() else delay)
delay = min(delay * 2, 30)
continue
r.raise_for_status()
data = r.json()
_cache[key] = (time.time(), data)
return data
raise RuntimeError("Still rate limited after retries")

def normalize(raw):
"""ADAPTER: map the provider response to our internal shape.

The raw shape below is an ASSUMPTION modeled on typical odds APIs.
Check the sandbox and edit only this function to match real fields.
"""
events = {}
for f in raw:
    books = {}
    for bm in f.get("bookmakers", []):
        for m in bm.get("markets", []):
            if m.get("key") != "h2h":
                continue
            prices = {o["name"]: float(o["price"])
                      for o in m.get("outcomes", []) if o.get("price")}
            if prices:
                books[bm["name"]] = prices
    if len(books) >= 2:
        events[str(f["fixture_id"])] = {
            "label": f"{f['home']} vs {f['away']}",
            "start": f.get("start_time", ""),
            "books": books,
        }
return events
Enter fullscreen mode Exit fullscreen mode

def fetch_odds(sport, mock=False):
if mock:
with open("mock_odds.json") as fh:
return normalize(json.load(fh))
key = os.getenv("ORBISTATS_API_KEY")
if not key:
raise SystemExit("Set ORBISTATS_API_KEY or run with --mock")
raw = _get(f"{BASE}/{sport}/odds", {"Authorization": f"Bearer {key}"})
return normalize(raw)

Rate limits matter more than they look. Polling 13 sports once a minute is 13 × 1440 = 18,720 requests a day. Check your plan, and if you need genuinely live prices, look at push delivery (the WebSocket API or webhooks) instead of polling. The HTTP details behind the retry logic are on MDN: 429 Too Many Requests and Retry-After.

Step 5: Mock Data So Anyone Can Run It

Save this as mock_odds.json. It has one match, a sharp reference book, two ordinary books, and one soft book with a generous home price:

json
[{
"fixture_id": 1001, "home": "Team A", "away": "Team B", "start_time": "2026-10-03T15:00:00Z",
"bookmakers": [
{"name": "sharpbook", "markets": [{"key": "h2h", "outcomes": [
{"name": "home", "price": 2.00}, {"name": "draw", "price": 3.60}, {"name": "away", "price": 3.90}]}]},
{"name": "bookA", "markets": [{"key": "h2h", "outcomes": [
{"name": "home", "price": 1.98}, {"name": "draw", "price": 3.55}, {"name": "away", "price": 3.85}]}]},
{"name": "bookB", "markets": [{"key": "h2h", "outcomes": [
{"name": "home", "price": 2.02}, {"name": "draw", "price": 3.50}, {"name": "away", "price": 3.90}]}]},
{"name": "softbook", "markets": [{"key": "h2h", "outcomes": [
{"name": "home", "price": 2.20}, {"name": "draw", "price": 3.40}, {"name": "away", "price": 3.50}]}]}
]
}]
Step 6: The CLI (main.py)
python

main.py

import argparse, os, time
import requests
from dotenv import load_dotenv
from ev import find_value
from fetch import fetch_odds

load_dotenv()
seen = set() # avoid alerting the same price twice

def notify(pick, stake):
msg = (f"+EV {pick.ev:.1%} | {pick.event} | {pick.selection} @ {pick.odds} "
f"on {pick.book} (ref: {pick.ref}) | stake ~{stake:.2f}")
url = os.getenv("DISCORD_WEBHOOK_URL")
if url:
requests.post(url, json={"content": msg}, timeout=10)

def run_once(args):
events = fetch_odds(args.sport, mock=args.mock)
picks = find_value(events, sharp_books=args.sharp, min_ev=args.min_ev)
if not picks:
print("No value found.")
return
print(f"{'EVENT':<20}{'BET':<8}{'BOOK':<11}{'ODDS':<7}{'FAIR%':<8}{'EV%':<8}{'STAKE'}")
for p in picks:
stake = args.bankroll * p.kelly
print(f"{p.event:<20}{p.selection:<8}{p.book:<11}{p.odds:<7.2f}"
f"{p.fair_prob*100:<8.1f}{p.ev*100:<8.1f}{stake:.2f}")
key = (p.event, p.book, p.selection, p.odds)
if key not in seen:
seen.add(key)
notify(p, stake)

if name == "main":
ap = argparse.ArgumentParser()
ap.add_argument("--sport", default="football")
ap.add_argument("--mock", action="store_true")
ap.add_argument("--min-ev", type=float, default=0.02)
ap.add_argument("--sharp", nargs="*", default=["sharpbook"])
ap.add_argument("--bankroll", type=float, default=1000)
ap.add_argument("--loop", type=int, default=0, help="repeat every N seconds")
args = ap.parse_args()
while True:
run_once(args)
if not args.loop:
break
time.sleep(args.loop)

Run it:

bash
python main.py --mock

You should see one row: softbook offering home at 2.20, flagged at roughly +7% EV against the sharp reference, with a small quarter-Kelly stake. The exact digits depend on the power de-vig, so run it and compare with your own math. Discord alerts use a standard webhook if you set DISCORD_WEBHOOK_URL.

Step 7: Point It at Real Data
Open the sandbox and copy a real odds response.
Edit normalize() until it produces the internal shape. This is usually a 10-minute job, and it's the only place that changes.
Make sure selection names are consistent across books (home/draw/away everywhere). A mismatch here silently ruins the de-vig.
Run python main.py --sport football and compare a few outputs by hand.
Try another sport. For two-way sports like tennis, basketball, or baseball (home/away only), the same code works because the de-vig handles any number of outcomes. The provider's docs describe one API shape across sports, but verify each sport's market names.
Step 8: Tests (test_ev.py)
python

test_ev.py

import math
from ev import devig, kelly_fraction, find_value

def test_devig_sums_to_one():
p = devig({"home": 2.0, "draw": 3.6, "away": 3.9})
assert math.isclose(sum(p.values()), 1.0, abs_tol=1e-6)

def test_devig_reduces_implied_probability():
p = devig({"home": 1.9, "away": 1.9})
assert math.isclose(p["home"], 0.5, abs_tol=1e-6) # symmetric market

def test_kelly_zero_when_no_edge():
assert kelly_fraction(2.0, 0.5) == 0.0

def test_finds_generous_price():
events = {"1": {"label": "A vs B", "start": "", "books": {
"sharp": {"home": 2.0, "away": 2.0},
"soft": {"home": 2.2, "away": 1.7}}}}
picks = find_value(events, sharp_books=["sharp"])
assert picks and picks[0].book == "soft" and picks[0].selection == "home"
bash
pytest -q
Step 9: Track Closing Line Value (The Honest Scorecard)

Winning or losing a few bets tells you almost nothing. A better signal is CLV: did you get a better price than the market's final one? If you consistently beat the closing fair probability, your process has an edge. Otherwise, short-term wins are luck.

python

tracker.py

import sqlite3

db = sqlite3.connect("picks.db")
db.execute("""CREATE TABLE IF NOT EXISTS picks(
id INTEGER PRIMARY KEY, event TEXT, book TEXT, selection TEXT,
odds REAL, fair_prob REAL, closing_fair_prob REAL)""")

def log_pick(p):
db.execute("INSERT INTO picks(event,book,selection,odds,fair_prob) VALUES(?,?,?,?,?)",
(p.event, p.book, p.selection, p.odds, p.fair_prob))
db.commit()

def clv(odds_taken, closing_fair_prob):
"""Expected value of the price you took, measured against the closing fair probability."""
return odds_taken * closing_fair_prob - 1.0

Just before kickoff, re-run the fair-probability calculation, save it as closing_fair_prob, and average clv() over dozens or hundreds of picks. For deeper testing on past data, see the historical data API.

What Goes Wrong in the Real World
Stale odds: the "value" is often a line the book already changed. Verify before you act, and treat huge edges as suspicious (that's what max_ev is for).
Different rules: dead-heat, void, and settlement rules differ between books, so the same-looking bet isn't always the same bet.
Limits and restrictions: bookmakers limit winning accounts. Your stake size and account longevity are real constraints.
Reference quality: a weak reference makes fake edges. Prefer a genuinely sharp book, and compare against closing lines.
Small samples: 100 bets prove almost nothing. Judge the process by CLV over a large sample.
Terms of use: check the data provider's terms on storing and redistributing odds, and never put your API key in client-side code.
Legality and risk: gambling law varies by country and state. This is an analysis tool, so make sure your use is legal where you are, and never risk money you can't afford to lose. If gambling stops being fun, contact a local support service.
Where to Take It Next
Add totals and spreads with a per-line matching key (the same line value across books).
Switch from polling to the WebSocket feed for live markets, and see Orbistats' 50-line live scoreboard tutorial for the relay and reconnect pattern.
Wrap the finder in FastAPI with a cache, using this Python + FastAPI tutorial, or build a typed client with the .NET version.
Learn more about the WebSockets API on MDN.
Add a backtest that replays historical odds and reports CLV, not just profit.
Recap

You built a pipeline that normalizes odds from any provider, removes margin, finds prices that beat a sharp reference, sizes stakes conservatively, and measures whether your process has real edge. The code is small on purpose, and the hard part is data quality, timing, and discipline.

Which sport or market should I add next? Tell me in the comments.

Top comments (0)