DEV Community

Yuhe He
Yuhe He

Posted on

Diffing Word Frequency in Public Investor Memos: Change Detection That Sells

Advisory documents get read in one pass: an investor memo, a due-diligence answer, a client strategy PDF. Nobody skims them twice. That makes them word-dense — every sentence is doing paid work, and small wording changes move decisions.

Which makes them a surprisingly clean data source. Collect a firm's memos across a cycle (quarterly letters, client notes, public decks) and diff successive versions. Not the ideas — the words. Three classes of change carry signal:

  1. Dropped hedges. "We believe the sector faces challenges" → "the sector faces challenges." Removing "we believe" is a conviction upgrade. Automated diffs flag exactly these: we believe, we think, in our view, expected, likely.
  2. New nouns. New proper nouns (companies, countries, instruments) appearing in consecutive documents mark attention shifts. A name that appears for the first time in a 20-page memo is a position being built.
  3. Frequency cliffs. A theme that drops from 14 mentions to 2 between versions is not balance — it's an exit described politely.

The engineering is small: PDF text extraction, sentence segmentation, a stoplist of hedge phrases, a mention frequency table, per-document diffs. No NLP model needed; a lexical diff plus a curated hedge dictionary captures most of the value, and word-frequency tables are auditable — you can show the client which sentence changed, which no embedding model beats for trust.

The privacy note: public memos only. This is the OSINT discipline — work where the data was published on purpose. No scraping behind logins, no leaked documents, just the trail firms leave deliberately.

What I learned shipping this as a product: buyers of intelligence don't pay for analysis quality they can't verify; they pay for change detection they can check. "I read the memo and felt a tone shift" sells nothing. "The word 'opportunistic' appeared 3 times in Q1, 0 in Q2, and here are the sentences" sells.

I packaged the hedge dictionary, diff pipeline, and report template in my Telegram & Web OSINT Bundle ($5) — the same collection discipline, generalized to public text streams. Free sample brief shows the output format.

The general rule: when a document stream is read densely by professionals, word-level deltas are the cheapest alpha in the room. Text is free; noticing is the product.

Runs free on GitHub Actions - no server, no paid APIs.

Top comments (0)