Most "opportunity radars" are dashboards that show you the same noisy feed everyone else
sees. I wanted the opposite: a small, boring pipeline that watches a handful of sources I
actually care about and only pings me when something matches. No SaaS bill, no paid APIs.
Here is the architecture, the parts that broke, and the rules I now follow.
The shape of it
Three stages, each replaceable:
- Collect — a set of independent watchers (RSS, a couple of public pages, an IMAP mailbox, a Telegram channel). Each watcher writes normalized items to one queue.
- Filter — rules that match on sender address (not display name) and on keywords.
- Notify — one message, deduplicated by a stable id, rate-limited.
The whole thing is a few hundred lines of Python plus Playwright for the JS-heavy pages.
Stage 1: collect, but keep watchers dumb
Each watcher does one thing and exits. If a watcher crashes, it must not take the pipeline
down. I run them from cron/systemd timers, not from a long-lived process that slowly leaks
memory. A cron job that fails is obvious; a daemon that wedges at 3am is not.
For email I use IMAP with a Gmail app password and readonly=True on select. Reading
should never be able to mutate the mailbox. For every message I compute a stable dedupe key
(Message-ID if present, otherwise a hash of from + subject + date) and skip anything seen.
def dedupe_key(msg):
mid = msg.get("Message-ID")
if mid:
return mid.strip()
raw = f'{msg.get("From","")}|{msg.get("Subject","")}|{msg.get("Date","")}'
return hashlib.sha256(raw.encode()).hexdigest()
Stage 2: filter on the address, not the name
This is the mistake I see in almost every filter I review. A rule like
if "Acme" in from_header matches "Acme Inc" <phish@evil.example>. Match the address:
from email.utils import parseaddr
addr = parseaddr(msg.get("From", ""))[1].lower()
Then match domains/substrings against that. It is a one-line change that closes an entire
class of false positives and phishing lookalikes.
Two more rules that pay for themselves:
-
Skip marketing. If the subject contains
newsletter,unsubscribe,promo,sale, it is not an opportunity. Drop it before it reaches the notifier. - Cap the noise. Never send more than N notifications per run. If the cap is hit, say so in the message. Silent truncation is worse than no notification.
Stage 3: notify once, and log
Every outbound notification gets an append-only on-disk log with a hash of the body, not
the body itself. If a credential ever ends up in an error path, the log must not become the
leak. Redact secrets in error strings before printing:
def redact(text, secrets):
for s in secrets:
if s:
text = text.replace(s, "<redacted>")
return text
What broke
-
Browser pages that render in JS.
requestsreturns an empty shell. Playwright with a persistent context fixed it, but you must target the active tab, not a new one, or you fight your own automation. - Probe noise in logs. A health check hammering an internal endpoint generated ~97% of the rows in my call log and drowned the signal. I added a prune command that only runs when explicitly invoked. Read-write by accident is how you lose data.
- Full test suites hanging. A suite that pulls in network tests will hang in CI. Run the targeted tests for the modules you touched; keep the slow integration ones behind a flag.
The rules I keep
- Collectors are interchangeable and disposable.
- Filter on addresses; treat display names as hostile.
- Deduplicate by a stable id, always.
- Redact secrets at the boundary; never log raw bodies.
- One notification, with its own audit line.
- Read-only by default; mutation requires an explicit command.
Total cost: one machine you already have, plus free-tier model calls. The value is not the
code — it is the discipline of saying no to 99% of the input so the 1% is visible.
I build small automation pipelines and write about them. If your team needs a watcher that
actually respects your attention, I can help — portfolio below.
Top comments (0)