DEV Community

Ramdai Bista
Ramdai Bista

Posted on Originally published at agentkitworks.com

Your Monitor, Scraper, and Database Are Three Scripts That Should Be One Pipeline

Most "watch a page and store what changes" setups start as three unrelated scripts: a cron job that polls for changes, a scraper that fetches the new content, and some code that writes it somewhere. Each one works in isolation. The bugs show up at the seams.

Why three scripts drift apart

Write the monitor, scraper, and storage step separately and you've created three places for the same fact to go stale independently. The monitor fires on any diff — including a timestamp in the footer or an ad slot rotating. The scraper re-fetches the whole page instead of just what changed. The storage step has no idea whether this run is a genuine update or a duplicate of one it already wrote an hour ago, so it either writes a dupe or — worse — someone adds a dedupe check that silently swallows a real change along with the noise.

None of these are hard problems individually. They're hard because nobody designed the handoff between the three stages, so each one makes its own private assumption about what the others already guaranteed.

The pattern: one pipeline, one shared key

The fix isn't a smarter scraper. It's treating "detect → fetch → persist" as one pipeline with a single identity key running through all three stages, instead of three jobs that happen to run in sequence.

  1. Detect on content hash, not raw diff. Hash the meaningful content (strip timestamps, ad blocks, session tokens) and compare hashes. A change in the hash is a real signal; a change in raw HTML usually isn't.
  2. Pass the hash forward as the fetch's dedupe key. The scrape stage checks "have I already stored this hash for this source?" before it does the expensive fetch — not after. This is the difference between catching a dupe for free and catching it after you've already burned a request.
  3. Make storage idempotent on (source_id, content_hash), not insert-only. An upsert keyed on that pair means replaying the same event twice is a no-op instead of a duplicate row. This one decision eliminates an entire category of "why do we have this record twice" bugs.
  4. Keep retries at the fetch boundary only. If the fetch fails, retry the fetch. Don't retry the whole pipeline from the monitor — that re-triggers detection logic that already succeeded and risks re-processing a hash you've moved past.
  5. Log the chain, not just the outcome. One pipeline run should produce one trace: detected hash X at time T, fetched Y bytes, stored under key Z. When something looks wrong three weeks later, you want to answer "what did this run actually do" without reconstructing it from three separate log files with no shared ID.

What this catches that three separate scripts don't

  • Orphaned fetches — the scraper ran and got data, but the storage step never got called because the two were wired by "run after" in a crontab rather than by passing the result forward. The data that was fetched is just gone, and nothing failed loudly.
  • Duplicate records from re-detection — the monitor fires again on a page that didn't meaningfully change (session token rotated), and because there's no shared dedupe key, storage happily writes a second copy.
  • Silent half-runs — a deploy or a timeout kills the process between fetch and store. With a shared pipeline and one trace per run, this shows up as an incomplete trace. With three independent scripts, it shows up as nothing, until someone notices the data's missing weeks later.

None of this requires fancy infrastructure — a queue or even a single function that calls all three steps in order, passing the hash and source ID through, gets you most of the value. We ended up wiring exactly this chain for ourselves because maintaining three independently-scheduled jobs for what is conceptually one pipeline kept producing exactly the bugs above (the Data Stack kit, if you'd rather not wire it yourself: https://agentkitworks.com/compare/data-stack-vs-separate).

The habit that actually matters, pre-wired or not: stop asking "did the monitor run, did the scraper run, did the write succeed" as three separate questions, and start asking "what did this one pipeline do, end to end, for this one key." That's the question three independent scripts can never answer, because nothing forces them to agree on what a single run even is.

Top comments (0)