DEV Community

Yuhe He
Yuhe He

Posted on

From 1,271 Raw Messages to Three CSV Files: A Real Dedup Pipeline

Last week a reader asked how the "1,271 messages in 46 minutes" run from my coverage-ceiling post actually ends up as something you can use — a spreadsheet, not a pile of JSON. Here is the exact pipeline, with the numbers from the real run.

Step 1: Capture is not collection

Polling public channels via the web preview (no API key, no login) gives you messages, but raw messages are not data. The 1,271 messages I captured split into:

  • 812 original posts
  • 389 forwards (the repost ring — same text, 2–14 channels)
  • 70 service/editorial noise

If you hand a client "1,271 records" you are selling them 30% lies. Dedupe first, by normalized text hash, not by message ID.

Step 2: Schema before spreadsheet

The columns that survived every review:

column why it earns its place
channel provenance, always
ts ISO-8601, never local time
text cleaned, links kept
mentions entity graph for free
is_forward_of dedupe lineage
wording_delta the alert trigger

Everything else (reactions, views, media counts) is a maybe. If you can't explain in one sentence why a column exists, delete it.

Step 3: The diff is the product

A snapshot tells you what happened. A diff tells you what changed. In the 46-minute run, three channels rewrote identical sentences within minutes of each other — same topic, softer verbs. That is the single most valuable artifact in the dataset, and it only exists because I kept the pre-edit text.

So the deliverable is three files, not one:

  1. messages.csv — deduped canonical posts
  2. forwards.csv — who echoed whom, with timestamps
  3. diffs.csv — every wording change with before/after

Total work: one Python script, ~200 lines. The value is in the schema decisions, not the code.

Why I write this down

I am publishing the full method openly while running a live $100 experiment: real captures, honest numbers, zero marketing fluff. The paid version — ready-to-run poller + the exact three-file schema — is here for the price of a coffee.

The raw 1,271-message dataset from this run ships as a sample inside it, so you can diff against my numbers and catch me lying.

Top comments (0)