A competitor content tracker is a scheduled job that fetches a list of competitor pages, compares each one against the copy it stored last time, and reports only the differences worth reading. Many teams build the naive version first. A cron entry curls the page, difflib compares the HTML, and your digest fills with changes nobody made. So the week a competitor cuts its price, that change looks like all the rest. This guide builds the version that stays quiet unless the change matters.
TL;DR
- A scheduled Python job stores the extracted body of each page and stays quiet while that body is unchanged.
- Decodo's web scraping solutions handle the fetch. They render JavaScript, can get past blocks that refuse a plain GET, and post each finished task ID to a callback URL when you supply one.
- Claude reads the diff and nothing else, then returns a change category plus a one-line summary. The category lets the digest route a launch differently from a nav tweak.
What you'll need
- Python 3.10 or newer, with requests, trafilatura, pydantic, and anthropic 1.8 or newer installed
- A Decodo account with Web Scraping API access, with its Basic header value in DECODO_TOKEN. This job spends 1 scrape per tracked page per check.
- An Anthropic API key in ANTHROPIC_API_KEY, on an account with credit
- A Slack incoming webhook URL in SLACK_WEBHOOK, or any URL that answers a POST with 200 while you test
- Any scheduler you already run, such as cron, a systemd timer, or GitHub Actions
Why checking competitors by hand stops working
Checking by hand is website change monitoring done with a bookmark folder and good intentions. It survives 2 competitors and falls apart as the list grows, because every check depends on a human remembering to look. A launch or a price cut gets spotted a week late, by someone who wasn't looking for it.
Live here means scheduled and near real time. For marketing pages, hourly is the tightest interval worth running, and the alert arrives when the scrape finishes. Nothing fires at the instant a competitor's CMS saves a page, and marketing content doesn't need it to fire.
Rule out the cheap paths first. If you only need to know whether a competitor published something, their /feed.xml or sitemap.xml with lastmod answers that in one small request. A conditional GET settles the unchanged case for the cost of a 304, anywhere If-None-Match is honored. This tracker is for pages with no feed, where the change is an edit rather than a new URL.
changedetection.io is the incumbent there. It self-hosts under Docker, and it ships change summaries of its own. Deploy it if you want a monitoring product. Build this if you want the signal inside a pipeline you already run, routed by rules you own.
Getting a change signal you can trust
Before the diff is worth reading, you have to handle 2 failure modes.
- The fetch fails without notice. I requested 4 pages with a browser User-Agent and no proxy, then requested the same 4 through Decodo:
| Target | Plain GET | Through Decodo |
|---|---|---|
g2.com/products/slack/reviews |
403 |
403 |
crunchbase.com/organization/stripe |
403 |
200 |
similarweb.com/website/stripe.com |
connect timeout |
200 |
vercel.com/pricing |
200 |
200 |
Only 1 of the 4 answered a plain request. Decodo recovered 2 of the 3 that refused, but G2 refused both attempts, which is where the fetch layer stops. The failure that hurts a tracker is the quiet one. A tracker that stores an empty body reports a huge change on the next good pull, and that false positive has a real fetch failure underneath it.
- The diff fires on nothing. I pulled each of 5 pages twice as raw HTML and twice as extracted Markdown, seconds apart, which is not long enough for anyone to have edited them. Raw HTML changed on 4 of the 5, by 2 lines at the low end and 10 at the high end, while all 5 extracted bodies came back unchanged. The noisiest page, and the 1 in 5 that never moved:
| Page | Changed lines, raw HTML | Changed lines, extracted |
|---|---|---|
vercel.com/blog |
10 |
0 |
posthog.com/blog |
0 |
0 |
None of that churn came from the copy. It was per-request identity every time, mostly tracing IDs and A/B assignment tokens regenerated on each render. Every extracted body ran 138 to 936 lines, so the zeros mean stable extraction and not an empty response. If you test on the 1 page in 5 that happens to be clean, raw HTML diffing looks fine until you add the other 4.
Which extractor you use decides how much of the remaining noise survives. I put one pair of captures of Linear's pricing page through both Decodo's Markdown conversion and through trafilatura, which drops navigation, footers, and sidebars:
| Extractor | Body | Changed lines | Of those, nav links |
|---|---|---|---|
markdown: true |
4,192 chars |
28 |
7 |
trafilatura |
1,943 chars |
19 |
0 |
A quarter of that first diff was navigation. Use markdown: true when you want no extra dependency, and a main-content extractor when the diff is what you act on. The 2026 WCXB benchmark scores the current field.
Decodo's web scraping solutions do the fetching half no extractor can do. With headless: "html", JavaScript renders before the markup reaches you, so a page that would arrive as an empty shell arrives as a real body. The geo parameter pins the request to a country, which matters when a competitor prices by region and when a rotating proxy pool would otherwise move the page under you. markdown: true is there when you would rather not add trafilatura.
A POST to /v3/task returns a task ID immediately. With callback_url set, Decodo posts to your endpoint at the moment the scrape finishes:
curl -X POST https://scraper-api.decodo.com/v3/task \
-H "Authorization: $DECODO_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url":"https://competitor.example/pricing","markdown":true,
"headless":"html","callback_url":"https://you.example/hook",
"passthrough":"pricing"}'
The callback carries the task ID, not the page body, so you still fetch the body from /v3/task/{task_id}/results, where it stays retrievable for 24 hours. passthrough comes back untouched and tells your handler which page landed. The script below polls instead, because polling needs no public URL.
You schedule the repeat runs yourself. Point cron, a systemd timer, or a GitHub Actions schedule at the POST that submits the task. Decodo's walkthrough on how to schedule scraping tasks covers the cron and GitHub Actions routes.
From diff to signal
Getting the page and reading the diff sit on opposite sides of a seam, and 3 outcomes never leave the fetch side:
The diff is clean enough to act on and still too long to read. To get a real change without waiting a month, I seeded the snapshot store from an archived copy of Linear's pricing page and ran the tracker against the live one. That run produced 19 changed lines, and the ones that matter sit among the rest. A 10-line excerpt:
-- Linear Agent (beta)
+- Linear Agent
-- Code Intelligence (beta)
+- Code Intelligence
-Trusted by more than 33,000 companies
+Trusted by more than 40,000 companies
+Coding sessions**
+Loops**
-Data warehouse sync
+** Requires AI credits
Only the footnote that puts 2 plan features behind AI credits changes what a sales team says on a call, and it lands after 18 lines that mostly shuffle beta labels. Claude sorts through those lines, reading the diff and never the page.
The 3 blocks below are one file, tracker.py, in that order. Half of it is the 2 API calls. Decodo returns the rendered page, Claude turns a diff into a fixed shape, and nothing else touches a network:
import difflib, hashlib, os, re, sqlite3, time, requests, trafilatura
from typing import Literal
import anthropic
from pydantic import BaseModel
DECODO = "https://scraper-api.decodo.com/v3/task"
HEADERS = {"Authorization": os.environ["DECODO_TOKEN"], # "Basic <base64 user:pass>"
"Content-Type": "application/json"}
VOLATILE = re.compile(r"^\s*[\d,.]+\s*$") # lines that are only a live counter
PRICING_PAGES = {"https://competitor.example/pricing"} # list it in the run loop too
ACTION = {"new_content": "notify", "pricing_change": "page",
"positioning": "notify", "cosmetic": "drop"}
class Change(BaseModel):
category: Literal["new_content", "pricing_change", "positioning", "cosmetic"]
summary: str
worth_paging: bool
PROMPT = """Below is a unified diff of one competitor page's extracted body,
previous against current. Name the single most important change as
new_content (a post or page that did not exist), pricing_change (a number,
tier, limit, or plan name moved), positioning (headline or product
description rewritten), or cosmetic (nav, link order, asset paths, counters).
Summarize it in 25 words or fewer. Set worth_paging only for a pricing_change
or a named product launch.
<diff>
{diff}
</diff>"""
def fetch(url):
"""Submit an async task, then poll for the rendered HTML."""
body = {"url": url, "headless": "html", "proxy_pool": "premium"}
r = requests.post(DECODO, json=body, headers=HEADERS, timeout=60)
r.raise_for_status() # 401 means DECODO_TOKEN lost its "Basic " prefix
task = r.json()["id"]
for _ in range(60):
r = requests.get(f"{DECODO}/{task}/results", headers=HEADERS, timeout=60)
if r.status_code == 200: # 204 means still running
res = r.json()["results"][0]
return res["status_code"], res["content"]
time.sleep(2)
return None, ""
def body_of(html):
"""Main content only. Nav, footer and sidebars never reach the diff."""
text = trafilatura.extract(html, include_links=True, include_tables=True) or ""
return [l.strip() for l in text.splitlines()
if l.strip() and not VOLATILE.match(l)]
def classify(diff):
r = anthropic.Anthropic().messages.parse(
model="claude-sonnet-5", max_tokens=4096,
output_config={"effort": "low"},
messages=[{"role": "user", "content": PROMPT.format(diff=diff[:12000])}],
output_format=Change,
)
return None if r.stop_reason == "refusal" else r.parsed_output
That status_code == 200 check is load-bearing, and I added it only after a crash. A pending task answers /v3/task/{id}/results with 204 No Content and an empty body, and requests counts any 2xx as success, so a loop written on r.ok calls .json() on nothing.
Volume and stakes decide the model, and Anthropic's models overview lists the tradeoffs.
The other half is storage and a branch. It stops a failed fetch, a first sighting, and an unchanged page from ever looking like a change:
def check(conn, url, slack):
status, html = fetch(url)
lines = body_of(html) if status == 200 else []
body = "\n".join(lines)
prev = conn.execute("SELECT body, sha FROM pages WHERE url=?", (url,)).fetchone()
# A body that vanished, or halved, is a fetch problem and not an edit.
if not lines or (prev and len(body) < len(prev[0]) // 2):
conn.execute("INSERT INTO failures VALUES(?,1) ON CONFLICT(url) "
"DO UPDATE SET n=n+1", (url,)); conn.commit()
return "fetch_failed" # no snapshot written, no diff
sha = hashlib.sha256(body.encode()).hexdigest()
keep = ("INSERT INTO pages VALUES(?,?,?) ON CONFLICT(url) "
"DO UPDATE SET body=?, sha=?", (url, body, sha, body, sha))
if prev is None:
conn.execute(*keep); conn.commit()
return "first_seen" # seed the store, stay silent
if prev[1] == sha:
return "unchanged" # stay silent
diff = "\n".join(difflib.unified_diff(
prev[0].splitlines(), lines, "previous", "current", lineterm="", n=1))
change = classify(diff)
if change is None:
return "dropped"
if url in PRICING_PAGES and repriced(prev[0], body):
change.category, change.worth_paging = "pricing_change", True
if ACTION[change.category] != "drop":
sent = requests.post(slack, json={"text": f"*{change.category}* on <{url}>\n"
f"{change.summary}"}, timeout=15)
sent.raise_for_status() # a dead webhook must not look like delivery
conn.execute(*keep); conn.commit() # advance only once the change is delivered
return change.category
Two runs against an unchanged page print first_seen and then unchanged, with nothing sent anywhere. A forced 403 and a nav-only shell both returned fetch_failed, and the stored snapshot survived both.
To reach the alert path without waiting for a competitor, edit the stored body in tracker.db and set that row's sha to any other value, then run the tracker once more. The comparison is hash against hash, so an edited body on its own changes nothing. What you edit becomes the previous side of the diff, so write the older price in and let the live fetch supply the new one.
A failed fetch and an unchanged page look identical from a distance. check returns fetch_failed before it reaches the snapshot, so a block never overwrites a good body and never counts as a change. The snapshot also advances only once the alert is delivered, so a dead webhook leaves the change to be found again next run instead of losing it. The half-length check is there because a nav-only shell extracts to a handful of characters and would otherwise read as a total rewrite.
Pointed at the seeded Linear diff, the classifier returned this:
category : new_content
summary : Linear Agent and Code Intelligence exit beta; new features 'Coding sessions' and 'Loops' added, requiring AI credits.
worth_paging: False
That is 19 diff lines reduced to 3 facts worth knowing. I ran the same classifier over 3 more archived-against-live diffs, and 1 is worth your attention.
A Supabase pricing diff carried 52 changed lines with a dollar figure on them, but the model called it positioning and declined to page. That reads like a miss, so I checked. All 33 prices on the page appear in both versions, and the compute table had been reformatted rather than repriced, so the model was right and a currency regex would have paged on nothing.
That is the limit of diffing text. Diff prose when you want to know what changed, and compare extracted structure when you need to know whether a specific number did:
class Tier(BaseModel):
name: str
price: str
class Tiers(BaseModel):
tiers: list[Tier]
TIERS = ("List every named plan or compute tier on this pricing page with its price "
"exactly as written. Omit anything that has no price.\n\n")
def repriced(previous, current):
"""Which named tiers changed price. Empty means no price moved."""
def tiers_of(body):
r = anthropic.Anthropic().messages.parse(
model="claude-sonnet-5", max_tokens=8000,
output_config={"effort": "low"},
messages=[{"role": "user", "content": TIERS + body[:20000]}],
output_format=Tiers,
)
return {t.name: t.price for t in r.parsed_output.tiers}
a, b = tiers_of(previous), tiers_of(current)
return {k: (a[k], b[k]) for k in a.keys() & b.keys() if a[k] != b[k]}
if __name__ == "__main__":
conn = sqlite3.connect("tracker.db")
conn.execute("CREATE TABLE IF NOT EXISTS pages(url TEXT PRIMARY KEY, body TEXT, sha TEXT)")
conn.execute("CREATE TABLE IF NOT EXISTS failures(url TEXT PRIMARY KEY, n INT)")
for url in ["https://competitor.example/pricing"]:
print(url, "->", check(conn, url, os.environ["SLACK_WEBHOOK"]))
On that Supabase pair, repriced returns nothing, which is right. Edit one tier from $110 to $129 in the stored body and it returns {'Large': ('$110', '$129')}, so the empty result above is a finding and not a silent failure. The PRICING_PAGES branch runs it only where a number is the point, and a moved price then overrides whatever the prose diff concluded. It costs 2 extra model calls on a page that changed, and it turns the question you least want to get wrong into a comparison.
3 ways to get this wrong
- Diffing raw HTML instead of extracted text. Extract first, normalize second, and hash what's left, so neither the tracing ids nor the counter lines reach the comparison.
- Scheduling faster than the content moves. An hourly check spends 168 requests a week on a blog that publishes weekly, when a weekly check would have caught the same post. Match the interval to the page, so pricing and changelog pages earn hourly checks, a blog index daily, and an about page weekly.
- Sending the whole page instead of the diff. Counted with count_tokens on one run, Linear's extracted body came to 1,031 input tokens against 590 for the diff, so sending the page cost 1.7 times as much to learn the same thing. That 1.7x grows with the page, because a blog index keeps getting longer while the number of new posts per check stays flat. The tier check is the exception, and it does send the whole body, because a diff cannot show you the price of a tier that did not move.
Final thoughts
Getting the page and reading the diff are 2 different problems. Decodo's web scraping solutions own the first, and Claude owns the second without ever seeing a page it doesn't need. Merging them puts a parser bug in the same channel as a repositioned headline. Tracking several competitors is a loop over a URL list, and logging each change's category alongside the snapshot turns the store into a record of where a competitor puts its effort. Point the tracker at the one page you keep meaning to check, seed the store on the first run, and let the second run tell you whether anything moved.

Top comments (0)