I use a lot of open data registries, and they all share one annoying property: you find
out something changed after it matters. A vendor gets added overnight, a grade shifts,
an offer disappears — and my notes are silently stale.
So I built a tiny tracker for one of them, and it turned out to be simpler than I
expected. The whole thing is a single Python file with zero dependencies — no pip
install, no virtualenv, no API key. I want to walk through the pattern because it
generalizes to any registry that publishes JSON.
The registry
Sourcey publishes three public datasets as plain JSON plus
a release feed: a companies registry (543 companies when I last checked), startup
credits, and an agent-readiness table with letter grades. On 2026-10-02 the live
release reported sha256:8094c042 consistently across all three datasets — that's
the number that makes verification possible.
The key detail: every dataset snapshot is addressable and the current release carries
a digest. That's all you need for real change detection — not "fetch and eyeball" but
"prove the bytes changed."
The tool
sourcey-tracker — one file, ~450 lines,
stdlib only. It gives you five commands:
$ python3 tracker.py status # release, digest, entity counts
$ python3 tracker.py snapshot # immutable local snapshot with digest
$ python3 tracker.py verify # re-hash the snapshot (detects tampering)
$ python3 tracker.py diff snap1 snap2 # field-level changes between releases
$ python3 tracker.py changelog snap1 snap2
$ python3 tracker.py watch --once # CI-friendly: exit 1 if anything changed
status gives you the dated-fact view of the registry — release id, when it was cut,
and the entity counts per dataset. diff is the interesting one: it walks both
snapshots and reports added entities, field-level changes, and removals with old and
new values. When I mutation-tested it — adding a fake company, changing one grade from
A to F, deleting an offer — it caught all three.
Why zero dependencies is a feature, not a flex
Every dependency is a supply-chain vote you cast on behalf of everyone who runs your
tool. urllib.request has been in the stdlib since Python 2.something, hashlib
ships with SHA-256 built in, and json is right there. For a read-only tracker the
web platform of 2008 is genuinely enough.
The practical payoff: anyone can audit the whole attack surface in one sitting. A
reader can verify the tool isn't exfiltrating anything by reading one file — no
pip-audit ritual, no lockfile archaeology. For security tooling that's not a nice
to have, it's the whole point.
I've been applying the same discipline to the rest of the mini-toolchain:
sourcey-registry-cli queries the
datasets from the command line,
readiness-badges renders the grades as
SVG badges, x402-balance watches a wallet
for USDC arriving on Base via a raw JSON-RPC eth_call, and
franticls scans a public bounty board with
price and claim-slot pressure. One file each, stdlib only, MIT licensed.
The verification pattern
The part I'd actually copy into other projects is the snapshot/verify split:
-
snapshotwrites{data, meta: {source, release_id, sha256, taken_at}}and prints the digest. -
verifyre-reads the file, re-computes the digest over the canonical JSON serialization, and compares. -
diffcompares two snapshots by digest first — identical digests short-circuit to "no changes" without walking the data.
Canonical serialization (sorted keys, fixed separators) matters more than people
think. If you hash json.dumps(obj) with default settings you'll get false diffs from
key ordering alone.
What I learned mutation-testing my own diff
I didn't trust diff until I broke it on purpose:
-
Added entity — new key in the companies map → reported under
added -
Changed field — one company's
region: "EU"→"US"→ old and new values printed -
Removed entity — deleted offer → reported under
removed -
Grade change —
A→F→ surfaced with the entity name and field path
The failure mode that surprised me: my first version compared entity sets but not
field values, so a grade change produced no output at all. The lesson — for a tracker,
"same entities" and "same data" are different claims, and you have to test both.
Use it on other registries
The pattern is deliberately registry-agnostic: fetch JSON → canonicalize → digest →
snapshot → diff. If a registry you care about publishes versioned JSON (and more of
them do, because agents need stable URLs to cite), you can point this shape at it in
an afternoon. The only registry-specific code in the tool is the endpoint table at
the top of the file.
Go break something on purpose first though. A tracker you haven't mutation-tested is
just a vibes engine.
All tools referenced are MIT licensed and single-file. If you track a public
registry with something better, I genuinely want to see it — the whole point of
doing this in the open.
Top comments (0)