DEV Community

Delta Tools
Delta Tools

Posted on

How I Built a Substack New-Post Monitor (and the Dedup Problem Nobody Warns You About)

TL;DR: Fetching a Substack publication's latest posts took an afternoon. Making the monitor not alert twice on the same post — across runs, crashes, and machine moves — was the actual project. Here's the naive version, the three ways it broke, and the dedup design that finally held. The finished tool is our Substack New-Post Monitor on the Apify Store.


The naive version (afternoon one)

The job sounded trivial: watch a Substack publication, tell me when a new post appears. Substack even hands you the data — every publication has an RSS feed at <publication>.substack.com/feed. Parse it, compare against last time, alert on the difference.

import feedparser

feed = feedparser.parse("https://example.substack.com/feed")
latest = feed.entries[0]
print(f"Latest: {latest.title} — {latest.link}")
Enter fullscreen mode Exit fullscreen mode

Done, right? It printed the latest post. I wired it to a Slack webhook, set a 15-minute cron, and felt extremely productive for about a day.

Then the double alerts started.

Break #1: "What did I already see?" has no obvious home

The first version kept seen_ids.json next to the script. It worked until I moved the script to a different machine and every post re-alerted — the state file didn't come along. Then I forgot the file on a deploy and got a fun burst of "NEW POST" messages for articles from 2023.

The lesson: the state file is the product, and it needs to live where the compute lives. On Apify, that meant a named key-value store shared across scheduled runs — state travels with the actor, not with my laptop. Anywhere else, it means a real database or at minimum a state file on the same always-on host as the cron job, backed up.

Break #2: The crash between "detect" and "save"

This one was subtler. The original loop was:

  1. Fetch feed
  2. Diff against seen IDs
  3. Send alerts for new posts
  4. Save updated seen IDs

If the process died between 3 and 4 — machine reboot, OOM kill, deploy — the next run re-detected the same posts and re-alerted. I "fixed" it by saving state before sending alerts, which flipped the failure mode: crash between save and alert meant silently missing posts, which is worse. A missed alert is invisible; a duplicate is just annoying.

The correct order is a tiny transaction: fetch → diff → persist the new seen-set first → send alerts → if alerting fails, the next run sees the posts as already-seen... no wait, that drops them again.

What actually works: persist seen IDs and a pending-alerts queue atomically, then drain the queue. On startup, drain any leftover queue before doing new work. It's a miniature outbox pattern, and yes, it's overkill for a side project — until the one time your 3am deploy eats the alert for the post you were actually waiting for.

For the Apify version, the platform's dataset + key-value store gives atomic-ish primitives that make this straightforward. The principle is the same anywhere: never let "I noticed it" and "I told someone" be separable without a recovery path.

Break #3: Identity is harder than it looks

What identifies "the same post" across polls? My first attempt used the post URL. Fine — until Substack's feed briefly served a post with a different canonical URL during an update, and it re-alerted. Then I tried titles; Substack authors edit titles, and every edit re-alerted.

RSS entries have GUIDs (entry.id in feedparser), which are supposed to be stable identifiers. They mostly are. My prototype scheme: GUID primary, URL as fallback, and never the title. Two identifiers agreeing beats one identifier you trust.

The shipped version simplified this further: Substack's archive API hands you stable post IDs, so the dedup key is just the post ID. Same principle, less machinery.

What the finished actor does

The Substack New-Post Monitor on the Apify Store is this design, packaged:

  • You give it a publication URL and a schedule. It fetches via Substack's public archive API, dedups on Substack's stable post IDs, persists its snapshot in a named key-value store (shared across scheduled runs — the lesson from Break #1, applied), and emits only genuinely new posts to the output dataset.
  • New posts land in the dataset ready to wire to alerts via Apify webhooks; a check that finds nothing costs a fraction of a cent.
  • It passed 7/7 of my stress tests: duplicate suppression across runs, crash recovery, multi-publication state isolation, feed hiccups.

I'm not claiming it's the only way to do this. The 20-line script from my alerts tutorial plus a seen_ids.json on a Raspberry Pi is a completely legitimate setup — I ran exactly that for weeks. The actor is for when you'd rather not be the ops team for your own alerts.

The real lesson

Every "monitor X for changes" tool is two tools: a fetcher (afternoon) and a rememberer (the actual project). The fetcher is the demo. The rememberer is the product. If you're building any kind of change-detection — price trackers, job alert bots, listing monitors — budget your time accordingly: 20% fetching, 80% state, identity, and crash recovery.

FAQ

Why not just use Substack's email subscriptions?
They work, but they're inbox-bound and all-or-nothing per publication. A monitor routes to Slack/Discord/webhooks and watches without subscribing.

Does it handle paid publications?
Only public/free posts via the public feed — same limitation as any RSS-based approach. Paywalled content needs an authenticated session, which is a different tool.

How much does it cost to run?
On Apify's pricing: a scheduled check that finds nothing costs a fraction of a cent. You pay meaningfully only when there's actually a new post to report. (Exact numbers depend on Apify's current compute pricing.)

Could I build this myself?
Absolutely — the tutorial version is 20 lines plus a state file. Build it if you enjoy the ops; use the actor if you don't.

Top comments (1)

Collapse
 
mythex profile image
Mythex •

The detour in Break #2 ("no wait, that drops them again") is the most honest part of the post. That's exactly how it goes in real life.

One gap the outbox still leaves: draining the queue is at-least-once. If the process dies after Slack accepts the message but before the item is marked sent, the next run posts it again. Two ways to narrow it:

  • Post through the API instead of an incoming webhook. chat.postMessage returns the message's ts, so store it with the queue item right away; on retry, if a ts is already stored, skip it or call chat.update instead of posting again. Discord webhooks can return the message ID too (with ?wait=true) and support editing by ID.
  • If you stay on webhooks, put the post ID in the message text, so a duplicate at least reads as the same alert rather than a second new post.

How does the actor treat a post that Substack republishes with a new date? Same ID, so silent, or an "updated" event?