You ship a job board, a scraper, or a "data for AI agents" API — and one day a feed starts returning 3 records instead of 300, or quietly drops a field, and nothing tells you until a customer notices. A 200 OK is not a health check.
I keep getting bitten by exactly this, so I built jobfeed-watchdog: a small, copy-paste GitHub Action that fails your build when a job/data feed silently breaks — schema change, record-count collapse, or outage. Stdlib-only, MIT, ~30 seconds to wire up.
The failure mode it catches
A feed "works" (returns 200) but:
- the JSON shape changes — a key you depend on disappears,
- the record count collapses (50 -> 5) because a board changed its API,
- the endpoint is down or returns an error body with a 200.
None of that trips a naive curl in CI. This is the quiet class of breakage that takes days to notice.
What it does
For each feed you list, it computes a lightweight structural fingerprint (container type, the set of record keys, the count) and compares it to a stored baseline. On drift, it prints a report and exits non-zero — so a normal GitHub Actions step turns red and can page Slack / email / Discord.
$ python3 watchfeed.py feeds.json ./state
[
{
"count": 50,
"id": "remote-jobs-api",
"ok": true,
"problems": [],
"url": "https://remote-jobs-api.tten.no/v1/jobs"
}
]
RESULT: ALL HEALTHY
That's the real output from a live run against a real feed. When something's wrong you get the specific problem, not "it's fine":
schema changed: missing keys ['salary_max']count collapse: 50 -> 3 (ratio 0.06 < 0.5)feed unreachable / non-2xx
The 30-second setup
git clone https://github.com/earnnova-dev/jobfeed-watchdog
cd jobfeed-watchdog
cp -r watchfeed.py .github/feeds.json /path/to/your/repo
Then list the feeds you care about in feeds.json:
{
"feeds": [
{
"id": "my-job-feed",
"url": "https://example.com/api/jobs",
"record_path": "results",
"count_ratio": 0.5,
"min_count": 5,
"timeout": 30,
"headers": { "Authorization": "Bearer ***" }
}
]
}
The included .github/workflows/jobfeed-watchdog.yml runs it on a schedule (and on push). Add auth headers per feed if needed. That's the whole integration.
Design choices
- Stdlib-only core (urllib). No heavy installs in CI. YAML support is optional (pyyaml) — JSON feeds need nothing extra.
- Deterministic, pure functions. No network in the unit tests — I test the fingerprinting logic directly.
- Operator recovery is built in. When a real schema change happens (a feed legitimately adds a field), you re-bless the baseline through the tool's own establish path rather than hand-editing state. I did this myself when the feed it ships with gained four structured-salary fields — the watchdog correctly held red, then I re-blessed and it went green and stayed idempotent.
Why I built it
I maintain a set of small, no-auth job-data products, and "the feed changed under me" is the #1 silent-failure class I keep triaging by hand. I wanted the check to be boring: a file in my repo that a CI job runs and that only goes red when a real contract broke. If you pull job/gig/data feeds and don't want to find out about breakage from a customer, this is the guard I wish I'd had earlier.
It's MIT — clone it, point it at your feeds, and let your build be the one that notices.
Top comments (0)