Everyone says every YC startup is now the same AI agent. I ran cosine similarity on all 6,142 of them to see whether that is true, and by how much. In 2015, one in five companies in a batch had a near twin among earlier YC companies. In 2025 it was two in five.
6,142 Y Combinator companies with a live domain, 48 batches from Summer 2005 through Fall 2026. Each one got a vector from text-embedding-3-large at 1024 dimensions.
The catch: I did not embed homepages. Homepage copy varies more by who wrote it than by what the company does, and an embedding of raw copy mostly measures writing style. Instead, gpt-5-mini wrote a five-sentence summary of each company from two inputs, a rendered homepage (ScrapingBee) and search snippets for the domain (Serper). The prompt forces the same five questions in the same order every time: product, buyer, delivery, market, pricing. Same shape of text for a 2008 company and a 2026 one. 206 companies had no usable homepage and fell back to YC's own description; those skew old, so they push against the result rather than for it.
Similarity is cosine between the normalised vectors.
Test 1: does a company have a twin among earlier YC companies?
For each company, take its best cosine score against a random 500 companies from earlier batches. Count how many in the batch clear 0.80. Repeat with ten different random 500s and average.
- 2013 to 2019 batches: 18.3% of companies have such a match
- Summer 2025: 46.5%
- Winter 2023 is the first batch above 30%
Full analysis and data here: https://fundingwatcher.com/research/yc-batch-similarity/
Top comments (1)
Your 46.5% vs 18.3% twin-rate is a striking result, but the denominator quietly biases it: each company is compared against a random 500 from earlier batches, so a Summer 2025 company draws candidates from roughly 6,000 predecessors while a 2013 company drew from only a few hundred — the 'has a twin >=0.72' rate mechanically rises with batch position even if underlying similarity were flat. Normalizing per-candidate — mean >=0.72 matches per comparison rather than per company — would separate genuine convergence from pool-size inflation. Worth a second test: Winter 2023 is the first batch above 30%, so re-running with a fixed-size predecessor window per batch would show how much of the curve survives once pool size is held constant.