Last week a reader asked how the "1,271 messages in 46 minutes" run from my coverage-ceiling post actually ends up as something you can use — a spreadsheet, not a pile of JSON. Here is the exact pipeline, with the numbers from the real run.
Step 1: Capture is not collection
Polling public channels via the web preview (no API key, no login) gives you messages, but raw messages are not data. The 1,271 messages I captured split into:
- 812 original posts
- 389 forwards (the repost ring — same text, 2–14 channels)
- 70 service/editorial noise
If you hand a client "1,271 records" you are selling them 30% lies. Dedupe first, by normalized text hash, not by message ID.
Step 2: Schema before spreadsheet
The columns that survived every review:
| column | why it earns its place |
|---|---|
channel |
provenance, always |
ts |
ISO-8601, never local time |
text |
cleaned, links kept |
mentions |
entity graph for free |
is_forward_of |
dedupe lineage |
wording_delta |
the alert trigger |
Everything else (reactions, views, media counts) is a maybe. If you can't explain in one sentence why a column exists, delete it.
Step 3: The diff is the product
A snapshot tells you what happened. A diff tells you what changed. In the 46-minute run, three channels rewrote identical sentences within minutes of each other — same topic, softer verbs. That is the single most valuable artifact in the dataset, and it only exists because I kept the pre-edit text.
So the deliverable is three files, not one:
-
messages.csv— deduped canonical posts -
forwards.csv— who echoed whom, with timestamps -
diffs.csv— every wording change with before/after
Total work: one Python script, ~200 lines. The value is in the schema decisions, not the code.
Why I write this down
I am publishing the full method openly while running a live $100 experiment: real captures, honest numbers, zero marketing fluff. The paid version — ready-to-run poller + the exact three-file schema — is here for the price of a coffee.
The raw 1,271-message dataset from this run ships as a sample inside it, so you can diff against my numbers and catch me lying.
Top comments (0)