DEV Community

Cover image for How I Built a Real-Time Arbitrage Betting Scanner with Python and an Odds API
orbistats
orbistats

Posted on

How I Built a Real-Time Arbitrage Betting Scanner with Python and an Odds API

Odds comparison looks simple from the outside: fetch prices, find the best one, done. In practice, the hard part is not the maths. It is getting clean, consistent, fast data from several bookmakers without writing a separate parser for each one.

In this tutorial, we will build a working arbitrage scanner in Python. It pulls odds for multiple sports, finds the best price for every outcome across bookmakers, checks whether the implied probabilities add up to less than 100%, calculates how to split a stake, and sends an alert. Then we will upgrade it from polling to real-time streaming.

We will use Orbistats for the data layer, because its Odds API returns normalized odds across bookmakers and markets in one consistent schema, and the same API shape works across all 13 supported sports.

A quick honesty note before we start: arbitrage is not risk-free money. Odds move between the moment you see them and the moment you place a bet, markets get suspended, bets get voided, and bookmakers limit accounts that win consistently. Treat this project as a data-engineering and maths exercise, and always check the terms and local laws that apply to you.

What we are building

The scanner has five parts:

An API client that authenticates and fetches odds
A normalizer that flattens the response into one predictable structure
An arbitrage detector that finds the best price per outcome
A stake calculator that splits your bankroll so profit is equal on every outcome
An alert sender, plus a WebSocket mode for live updates

By the end you will have one file of roughly 200 lines that you can extend into a proper trading dashboard.

Prerequisites

You need Python 3.10 or newer and a free API key. You can create a free Orbistats key without a card, and if you want to see raw responses before writing any code, open the public API Sandbox.

bash
mkdir arb-scanner && cd arb-scanner
python -m venv .venv && source .venv/bin/activate
pip install requests python-dotenv websockets

Create a .env file:

bash
ORBISTATS_API_KEY=your_key_here
ORBISTATS_WS_URL=wss://... # copy the exact URL from the WebSocket docs
TELEGRAM_BOT_TOKEN=optional
TELEGRAM_CHAT_ID=optional

The API uses Bearer authentication and a versioned base URL, https://api.orbistats.com/v1/. The documentation covers the sport → resource → endpoint model, and the developer hub links to the SDK examples and status page.

Step 1: The arbitrage maths in 60 seconds

Every decimal price implies a probability: 1 / odds. If you take the best available price for each outcome of an event and the implied probabilities add up to less than 1, you can back every outcome and lock in a profit regardless of the result.

Take a tennis match with two outcomes:

Player A best price: 2.10 at Bookmaker X (implied 47.62%)
Player B best price: 2.05 at Bookmaker Y (implied 48.78%)

The total is 96.40%, so the margin is 1 / 0.9640 - 1, about 3.7% before fees and limits. If the total is 100% or higher, there is no arbitrage.

For a three-way market such as football 1X2, the same rule applies with three outcomes.

Step 2: The API client

Polling APIs need timeouts and retries, otherwise one slow response freezes your scanner. We will use a requests.Session with a retry adapter.

python

scanner.py

import os
import time
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
from dotenv import load_dotenv

load_dotenv()

API_KEY = os.getenv("ORBISTATS_API_KEY")
BASE_URL = "https://api.orbistats.com/v1"

13 supported sports: use the slugs from the documentation.

SPORTS = [
"football", "basketball", "american-football", "cricket", "tennis",
"horse-racing", "baseball", "esports", "golf", "combat-sports",
# add the 3 newest sport slugs here
]

def build_session() -> requests.Session:
session = requests.Session()
session.headers.update({"Authorization": f"Bearer {API_KEY}"})
retries = Retry(
total=3,
backoff_factor=0.5,
status_forcelist=[429, 500, 502, 503, 504],
)
session.mount("https://", HTTPAdapter(max_retries=retries))
return session

session = build_session()

def fetch_odds(sport: str) -> list[dict]:
"""Fetch current odds for one sport."""
res = session.get(f"{BASE_URL}/{sport}/odds", timeout=10)
res.raise_for_status()
return res.json().get("data", [])

Because the API keeps the same shape across sports, looping over all of them is just a for loop. That consistency is the main reason to use a normalized feed: with raw bookmaker feeds you would maintain a separate integration for each source, and one format change would break your scanner overnight.

Step 3: Normalize the response

Whatever the API returns, your own code should work with a small internal structure. That way, if a field name changes, you fix one function instead of your whole codebase.

Check the exact fields in the sandbox and adjust the mapping below. The important idea is to flatten everything into rows of event, market, outcome, bookmaker, price.

python
from dataclasses import dataclass

@dataclass(frozen=True)
class Quote:
event_id: str
event_name: str
market: str
outcome: str
bookmaker: str
price: float # decimal odds

def normalize_event(event: dict) -> list[Quote]:
quotes = []
name = f'{event["home"]["name"]} vs {event["away"]["name"]}'
for market in event.get("markets", []):
for line in market.get("outcomes", []):
price = line.get("price")
if not price or price <= 1.0:
continue # skip suspended or invalid prices
quotes.append(Quote(
event_id=str(event["id"]),
event_name=name,
market=market["key"],
outcome=line["name"],
bookmaker=line["bookmaker"],
price=float(price),
))
return quotes

Skipping suspended and invalid prices matters more than it looks. A stale or suspended line is the most common source of false arbitrage signals.

Step 4: Detect arbitrage opportunities

Group quotes by event and market, keep the best price for each outcome, and check the sum of implied probabilities.

python
from collections import defaultdict

def find_arbs(quotes: list[Quote], min_margin: float = 0.005):
grouped = defaultdict(dict) # (event, market) -> outcome -> best Quote

for q in quotes:
    key = (q.event_id, q.market)
    best = grouped[key].get(q.outcome)
    if best is None or q.price > best.price:
        grouped[key][q.outcome] = q

opportunities = []
for (event_id, market), outcomes in grouped.items():
    if len(outcomes) < 2:
        continue
    # only combine outcomes from at least 2 different bookmakers
    if len({q.bookmaker for q in outcomes.values()}) < 2:
        continue

    implied = sum(1 / q.price for q in outcomes.values())
    if implied < 1 - min_margin:
        opportunities.append({
            "event": next(iter(outcomes.values())).event_name,
            "market": market,
            "margin": 1 / implied - 1,
            "legs": list(outcomes.values()),
            "implied": implied,
        })
return sorted(opportunities, key=lambda o: o["margin"], reverse=True)
Enter fullscreen mode Exit fullscreen mode

The min_margin threshold filters out tiny 0.1% "opportunities" that disappear after rounding, fees and price movement. In practice, a floor of 0.5% to 1% keeps the noise manageable.

Step 5: Calculate the stakes

To guarantee the same return whichever outcome wins, each stake is proportional to its implied probability:

stake_i = bankroll × (1 / odds_i) / sum(1 / odds)

python
def calculate_stakes(opportunity: dict, bankroll: float = 100.0) -> list[dict]:
implied = opportunity["implied"]
plan = []
for leg in opportunity["legs"]:
stake = bankroll * (1 / leg.price) / implied
plan.append({
"outcome": leg.outcome,
"bookmaker": leg.bookmaker,
"odds": leg.price,
"stake": round(stake, 2),
"return": round(stake * leg.price, 2),
})
return plan

With a 100 unit bankroll and a 3.7% margin, every outcome returns about 103.70. Remember that bookmakers round stakes and enforce minimums, so the real profit is usually slightly lower.

Step 6: Put it together with a polling loop
python
def scan_once(bankroll: float = 100.0):
for sport in SPORTS:
try:
events = fetch_odds(sport)
except requests.RequestException as err:
print(f"[{sport}] fetch failed: {err}")
continue

    quotes = [q for e in events for q in normalize_event(e)]
    for opp in find_arbs(quotes):
        plan = calculate_stakes(opp, bankroll)
        print(f"\n[{sport}] {opp['event']} ({opp['market']}) "
              f"+{opp['margin']:.2%}")
        for leg in plan:
            print(f"  {leg['outcome']:<12} {leg['bookmaker']:<14} "
                  f"@{leg['odds']:.2f}  stake {leg['stake']}")
Enter fullscreen mode Exit fullscreen mode

if name == "main":
while True:
scan_once()
time.sleep(20) # respect your plan's rate limits

Polling every 15 to 30 seconds is fine for a side project and keeps you comfortably inside free-tier limits. But arbitrage windows are often short, and if you want a serious scanner, polling becomes the bottleneck.

Step 7: Go real-time with WebSocket

Instead of asking "has anything changed?" every 20 seconds, keep a persistent connection open and let the server push each price change. Orbistats provides a WebSocket API for exactly this, and the Live Scores API page explains how live match state fits in.

python
import asyncio, json, os
import websockets

WS_URL = os.getenv("ORBISTATS_WS_URL")
book = {} # (event_id, market, outcome, bookmaker) -> Quote

async def stream():
while True: # reconnect loop
try:
async with websockets.connect(
WS_URL,
additional_headers={"Authorization": f"Bearer {API_KEY}"},
) as ws:
print("Live feed connected")
async for message in ws:
update = json.loads(message)
for q in normalize_event(update):
book[(q.event_id, q.market, q.outcome, q.bookmaker)] = q
check_event(update["id"])
except Exception as err:
print(f"Disconnected ({err}), retrying in 3s...")
await asyncio.sleep(3)

def check_event(event_id: str):
quotes = [q for k, q in book.items() if k[0] == event_id]
for opp in find_arbs(quotes):
print("ARB:", opp["event"], f"{opp['margin']:.2%}")

asyncio.run(stream())

Two details are easy to forget. First, the reconnect loop: sockets drop on network hiccups and server restarts, and a scanner that silently goes stale is worse than one that never started. Second, keep an in-memory book of the latest price per bookmaker and re-check only the event that just changed, which keeps CPU usage flat even with thousands of live markets.

If you prefer server-side push without holding a socket open, the Webhooks API can call your endpoint when odds change. There is also a good Go webhook receiver walkthrough on the Orbistats dev.to profile.

Step 8: Send yourself an alert
python
def send_telegram(text: str):
token = os.getenv("TELEGRAM_BOT_TOKEN")
chat_id = os.getenv("TELEGRAM_CHAT_ID")
if not token or not chat_id:
return
requests.post(
f"https://api.telegram.org/bot{token}/sendMessage",
json={"chat_id": chat_id, "text": text},
timeout=10,
)

Call it inside your loop with the event name, margin and stake plan. Add a small cache of alerts already sent so you do not get the same message every 20 seconds.

Making it production-ready

Once the basics work, these upgrades matter most:

Freshness check: Ignore any quote older than a few seconds. Stale prices create most false positives.
Deduplication: Store a hash of each opportunity and alert only when it first appears or the margin changes meaningfully.
Backtesting: Use the Historical Sports Data API to replay past odds and measure how often your signals would have appeared and how long they lasted. This is where you find out whether a strategy is real or just noise.
Enrich with context: Combine odds with fixtures, standings and team form from the Sports Data API and the Sports Statistics API, so you can filter out events with lineup uncertainty or postponement risk.
Watch API changes: Versioning is documented, and breaking changes are meant to land behind new version numbers. Keep an eye on the changelog.
Add a UI: Instead of a terminal, build a dashboard. Orbistats also ships embeddable widgets, and the live scoreboard tutorial shows the WebSocket pattern in vanilla JS.
Common mistakes

Trusting every arbitrage signal. Most signals under 1% are stale prices, not opportunities. Add a freshness filter before anything else.

Ignoring market definitions. "Over 2.5 goals" at one bookmaker and "Over 2.5 goals (incl. extra time)" at another are different bets. Normalized markets reduce this problem, but always check the market key before combining legs.

Forgetting execution risk. Your scanner sees prices in milliseconds, but placing two bets takes seconds. Cap your stake size and assume some legs will fail.

Hammering the API. Polling faster does not help once you are at your rate limit. Move to streaming instead.

Where to go next

You now have a scanner that fetches normalized odds, finds cross-bookmaker gaps, sizes the stakes and streams live updates. From here you could add value-bet detection, line-movement tracking, or a FastAPI wrapper, like the one in the Python and FastAPI tutorial, so other services can query your signals.

Useful links to keep open while you build:

Orbistats homepage
Pricing and free tier
Full documentation
API Sandbox
About Orbistats
.NET Core version of the sports data tutorial

If you build something on top of this, share it in the comments. I would especially like to hear how you handle stale-price filtering and how much of a margin you consider worth alerting on.

Disclaimer: This article is for educational purposes only and is not financial or betting advice. Betting involves risk, may be restricted in your jurisdiction, and you should only stake what you can afford to lose.

Top comments (0)