DEV Community

Jakob Greenfeld
Jakob Greenfeld

Posted on

I ran cosine similarity on every YC startup. 2025 batches are 2.5x more repetitive than 2015

Everyone says every YC startup is now the same AI agent. I ran cosine similarity on all 6,142 of them to see whether that is true, and by how much. In 2015, one in five companies in a batch had a near twin among earlier YC companies. In 2025 it was two in five.

6,142 Y Combinator companies with a live domain, 48 batches from Summer 2005 through Fall 2026. Each one got a vector from text-embedding-3-large at 1024 dimensions.

The catch: I did not embed homepages. Homepage copy varies more by who wrote it than by what the company does, and an embedding of raw copy mostly measures writing style. Instead, gpt-5-mini wrote a five-sentence summary of each company from two inputs, a rendered homepage (ScrapingBee) and search snippets for the domain (Serper). The prompt forces the same five questions in the same order every time: product, buyer, delivery, market, pricing. Same shape of text for a 2008 company and a 2026 one. 206 companies had no usable homepage and fell back to YC's own description; those skew old, so they push against the result rather than for it.

Similarity is cosine between the normalised vectors.

Test 1: does a company have a twin among earlier YC companies?

For each company, take its best cosine score against a random 500 companies from earlier batches. Count how many in the batch clear 0.80. Repeat with ten different random 500s and average.

  • 2013 to 2019 batches: 18.3% of companies have such a match
  • Summer 2025: 46.5%
  • Winter 2023 is the first batch above 30%

Full analysis and data here: https://fundingwatcher.com/research/yc-batch-similarity/

Top comments (1)

Collapse
 
aifrontierpost profile image
AI Frontier Post •

Your 46.5% vs 18.3% twin-rate is a striking result, but the denominator quietly biases it: each company is compared against a random 500 from earlier batches, so a Summer 2025 company draws candidates from roughly 6,000 predecessors while a 2013 company drew from only a few hundred — the 'has a twin >=0.72' rate mechanically rises with batch position even if underlying similarity were flat. Normalizing per-candidate — mean >=0.72 matches per comparison rather than per company — would separate genuine convergence from pool-size inflation. Worth a second test: Winter 2023 is the first batch above 30%, so re-running with a fixed-size predecessor window per batch would show how much of the curve survives once pool size is held constant.