Every sports app starts the same way: you need scores, you don't have a budget, and some site out there already shows the data you want. So you write a scraper. It works. For about three weeks.
This post walks through why that approach breaks down, what it actually costs once you count your own time, and how I migrated a scraper-based sports app to a real API without a full rewrite — with code for both sides so you can see exactly where the pain is.
The Scraper Phase: It Works Until It Doesn't
Here's roughly what a first-pass scraper looks like — nothing fancy, just requests and BeautifulSoup pulling scores off a public scoreboard page:
python
import requests
from bs4 import BeautifulSoup
def get_live_scores():
resp = requests.get(
"https://example-sports-site.com/live-scores",
headers={"User-Agent": "Mozilla/5.0"},
)
soup = BeautifulSoup(resp.text, "html.parser")
matches = []
for card in soup.select(".match-card"):
matches.append({
"home": card.select_one(".home-team").text.strip(),
"away": card.select_one(".away-team").text.strip(),
"score": card.select_one(".score").text.strip(),
})
return matches
Twenty lines, no signup, no cost. For a weekend project, this feels like a win. Then reality shows up in this order:
Week 1 — it just works. You demo it, everyone's happy.
Week 2 — the site redesigns a div. .score becomes .score-value, and your scraper silently returns empty lists instead of erroring, which is worse than crashing because nobody notices until a user complains.
Week 3 — rate limiting or a CAPTCHA shows up. As one detailed scraping-vs-API breakdown puts it, many sites employ anti-scraping measures like CAPTCHAs or rate limiting, and layout changes routinely break scraping code that isn't built to expect them.
Week 4 — you start rotating IPs and user agents just to stay online, which is exactly the point where a hobby project turns into an ops job.
What Scraping Actually Costs (It's Not Free)
The part that got me was realizing "free" scraping wasn't free at all, once I counted my own hours. A detailed cost breakdown from BounceWatch's API-vs-scraping analysis makes the same point for a different data category: the biggest mistake teams make is comparing an API subscription price to zero, when scraping is never actually free — a production-grade scraper that handles errors, retries, rate limits and data normalization takes 40 to 80 hours to build properly, and any serious scraping operation at scale typically needs rotating residential proxies, which run $200 to $1,000 a month on their own.
For a live sports feed specifically, add: rebuilding selectors every time the source site redesigns, monitoring for silent failures (empty results are worse than errors), and the fact that scraping infrastructure never gets the JSON schema stability an API gives you. Scrape.do's comparison guide frames the core difference well: an API rate limit is a hard ceiling you can plan around, while a scraping rate limit is an engineering problem you have to keep re-solving.
The Legal Question Nobody Wants to Answer
This is the part most scraper tutorials skip entirely, and it matters more for a sports/betting-adjacent app than almost any other category.
The short version, pulled from a 2026 compliance guide across seven countries: the legal lines generally run through four boundaries — don't bypass authentication, be careful with personal data (GDPR/CCPA), extract facts rather than creative expression (copyright), and don't hammer a server hard enough to cause harm (rate limiting). The same guide notes something specific and important: registering an account, agreeing to terms that ban automation, and then automating anyway is the exact fact pattern that loses breach-of-contract cases. If the sports site you're scraping requires any kind of login or has a "no automated access" clause in its terms, you've crossed from a gray area into a contract you explicitly agreed to violate.
ScrapingBee's comparison is blunter about it: scraping publicly visible data is typically fine, but you must respect robots.txt, terms of service, and data privacy laws — and for anything commercial, get real legal counsel rather than assuming a blog post settles it. A separate ranking from Oxylabs puts it plainly too: public APIs carry the lowest legal risk of the available options, with custom scraping carrying the most.
For a hobby project displaying scores to friends, this risk is mostly academic. For anything you're putting a business or a user base behind, it stops being academic fast.
What Actually Changed When I Switched to an API
Here's the same "get live scores" function, this time against a real sports data API instead of a scraped page:
python
import os
import requests
API_KEY = os.getenv("SPORTS_API_KEY")
BASE_URL = "https://api.orbistats.com/v1"
def get_live_scores():
resp = requests.get(
f"{BASE_URL}/football/live",
headers={"Authorization": f"Bearer {API_KEY}"},
timeout=10,
)
resp.raise_for_status()
return resp.json()
Roughly the same line count, but almost everything about how it behaves is different:
Scraper API
Breaks when the source redesigns its page Yes, silently No — versioned schema
Rate limits Undocumented, discovered by trial and error Published, plannable
Auth None, or fragile session cookies Standard Bearer token
Error handling Empty results look like success HTTP status codes tell you what happened
Legal standing Depends on ToS and jurisdiction Covered by the provider's terms
Live/real-time data Requires constant polling of a page not built for it WebSocket/webhook delivery available
That "versioned schema" row is the one that saved the most time in practice. A REST-vs-scraping breakdown notes that APIs return deterministic, consistent data structures on every request — the same endpoint returns the same JSON shape, which makes writing robust parsing logic and catching anomalies straightforward. Scraped data, by contrast, varies based on user agent, geography, authentication state, A/B test variants and JavaScript rendering — none of which you control.
Migrating Without a Full Rewrite
The trick that made this manageable was not touching my application code at all — only the data-fetching layer. I already had a thin interface between "get scores" and "render scores," so the swap was one file:
python
data_source.py — this is the ONLY file that changed
class ScoreSource:
def get_live_scores(self):
raise NotImplementedError
class ScraperSource(ScoreSource):
def get_live_scores(self):
# old scraping logic here, kept for comparison/testing
...
class OrbistatsSource(ScoreSource):
def init(self, api_key):
self.api_key = api_key
self.base_url = "https://api.orbistats.com/v1"
def get_live_scores(self):
resp = requests.get(
f"{self.base_url}/football/live",
headers={"Authorization": f"Bearer {self.api_key}"},
timeout=10,
)
resp.raise_for_status()
return self._normalize(resp.json())
def _normalize(self, raw):
return [
{"home": m["home"]["name"], "away": m["away"]["name"],
"score": f"{m['home']['score']}-{m['away']['score']}"}
for m in raw
]
python
app.py — unchanged
source = OrbistatsSource(api_key=os.getenv("SPORTS_API_KEY"))
scores = source.get_live_scores()
render(scores)
Everything downstream of get_live_scores() — templates, WebSocket broadcasting, caching — never had to change. That abstraction layer is worth building on day one, even for a scraper, because it's exactly what makes a future migration a config change instead of a rewrite. I go into this pattern in more depth in my earlier post on picking a provider that won't force a migration later.
Live Data: Where the Gap Gets Bigger
Scraping a live scoreboard means polling the same page over and over, hoping the HTML hasn't changed and you haven't been rate-limited yet:
python
import time
while True:
scores = get_live_scores() # full page fetch + full HTML parse, every time
update_ui(scores)
time.sleep(15) # too slow for "live," too fast to be polite
A real API gives you push delivery instead. Using Orbistats' WebSocket API as an example:
python
import asyncio, json, os
import websockets
async def stream_scores():
async with websockets.connect(
"wss://YOUR_STREAM_HOST/PATH_FROM_DOCS", # copy exact URL from the WebSocket docs
additional_headers={"Authorization": f"Bearer {os.getenv('SPORTS_API_KEY')}"},
) as ws:
async for message in ws:
update_ui(json.loads(message))
asyncio.run(stream_scores())
No polling interval to tune, no full-page re-fetch every 15 seconds, and updates arrive the moment something changes rather than up to 15 seconds late. If you want the fuller build — REST for initial state, WebSocket for live deltas, reconnection handling — I covered that in detail in Build a Live Sports Scoreboard in 50 Lines of JavaScript with a WebSocket API.
Try Before You Commit
You don't need an API key to see if this is worth it. Orbistats' public sandbox lets you inspect real JSON responses with no signup, and a free tier is available on the pricing page if you want to test against your own code (free-tier live data has historically been delayed roughly 30–60 seconds, and plan details change, so check the live page). Full endpoint docs are on the Sports Data API page, and the documentation hub covers auth and every resource. Sport count on the platform has recently expanded (Orbistats now advertises 13 sports as of this writing — confirm the current list before building around a specific league), and the API shape is described as consistent across sports, so /football/ and /basketball/ follow the same pattern. If you're on a different stack, I've also written this exact "consuming the API" pattern in Go, .NET/C# and Python with FastAPI.
When Scraping Is Still the Right Call
To be fair to scraping: it's not always the wrong choice. Scrape.do's guide makes a point worth repeating — most of the web has no API at all, so if you need data that genuinely isn't offered through any official channel, scraping isn't one option among several, it's the only one. Scrape when the API you'd otherwise use leaves out data you actually need, when its rate limits genuinely won't let you finish a one-off research task, or when no API exists for that specific source. What doesn't hold up is scraping a data category — like sports scores and odds — where mature, documented, rate-published APIs already exist, purely to avoid a subscription fee.
Checklist: Should You Scrape or Use an API?
Does an API already exist for this data category? If yes, start there.
Will this run in production, or is it a one-off research script? Production leans API.
Do you need live/real-time updates? Scraping a page on a timer is a worse version of what push delivery (WebSocket/webhooks) already solves.
Have you read the source site's terms of service and robots.txt? If automation is explicitly prohibited and you'd need to log in to get the data, that's a real legal signal, not a technicality.
Have you actually costed out proxy/IP-rotation infrastructure against a paid API tier? Scraping at scale is rarely cheaper once you count it honestly.
Is your data-fetching code behind an interface/abstraction layer? If not, do that first regardless of which approach you pick — it's what makes switching later cheap.
Wrapping Up
A scraper is a great way to validate an idea over a weekend. It's a bad long-term foundation for anything that needs to stay online, stay legal, or go live in real time. The switch cost me about a day of work, almost entirely in that one data-source file, because I'd kept the fetching logic isolated from the rest of the app from the start. If you're still in the scraper phase, that's the one change worth making today — not the migration itself, but the abstraction that makes the migration boring whenever you decide to make it.
Have you made this switch yourself? What broke first — the selectors, the rate limits, or the legal nerves? Let me know in the comments.
Top comments (0)