Every pricing page, changelog and "about us" page is a moving target. Most change-detection SaaS charges $15-50 per month to do what a 40-line script and a cron job do for free. I built mine on a Saturday afternoon; here is the design, including the traps that make naive versions spam you and then go silent.
The pipeline
Four stages: fetch, normalize, diff, alert. Every bug I ever had lives in one of them.
1. Fetch. curl with a browser user-agent, a 15-second timeout, and two retries. Store the raw HTML with a timestamp before you touch it. Raw-before-normalize matters: when the monitor breaks at 3am you want the bytes it saw, not the bytes your parser salvaged.
2. Normalize. This is where free beats paid. A naive full-page hash fires on ads, campaign timestamps, session ids, and A/B test flicker - within a week you are ignoring every email. So strip scripts, styles, and dynamic ids; extract the main content region; collapse whitespace; keep visible text lines only. Hash the normalized text, never the raw page.
3. Diff at two levels. Line-level diff tells you a page moved. Word-level diff inside a changed block tells you what it says now. The alert that pays you is word-level: "the 30-day refund clause became 14 days", not "bytes changed". I alert on sentences that changed, not documents that arrived - same rule I use on security advisories and investor memos, and it is the rule that turns a spam feed into a signal feed.
4. Alert with a threshold. Fire only when the changed region exceeds N words (I start at 15). Below that is CDN noise. Above it, send yourself the before/after word diff, not just a URL - an alert you can read without opening the site is an alert you will actually keep reading.
The trap that kills most hobby monitors
A monitor that dies quietly is worse than no monitor. Bot walls go up, selectors rot, the site moves to a framework your extractor eats alive. So the pipeline needs a liveness layer: after each run, record the extraction rate (extracted chars / page bytes). A 403 spike, a zero-text response, or extraction collapse is itself an event. Page-changed and monitor-broken are different pages of the same alert inbox; most free tools conflate them and lose your trust in the same week.
What paid tools are actually selling you
Managed residential proxies, JS-rendered screenshots, mobile views, team seats. If you need those, buy them - $19/month is cheap against a missed price change on a product line. If you need "watch 12 competitor pages and tell me in words what changed", the cron job wins.
The discipline that makes it sellable
Snapshots are the asset. An alert is one day old and worthless; a month of normalized snapshots is a corpus you can diff, chart, and sell. When a supplier rewrites terms, the thing your client pays for is the series of rewrites, not the last one. My collection playbooks - the normalization rules, the liveness thresholds, the wording-diff alert patterns, with the coverage-cursor notes for channels and pages alike - are in the bundle (https://heyuhe.gumroad.com/l/poddr), and the free sample brief (pay what you want, even $0) is https://heyuhe.gumroad.com/l/ruldgm so you can check the density before you pay anything. Mail to heyuhe2003@gmail.com if you want a custom 48h report on your five watched pages.
The whole stack is curl, a text extractor, difflib, cron, and email. Twenty minutes to build, twenty seconds a day to read. The vendor's real product is the discipline in stage 2 and the honesty in stage 4 - and that part you can copy.
Top comments (0)