DEV Community

Cover image for There Are Now Two Internets. You're Only Optimizing for One.
Phil Rentier Digital
Phil Rentier Digital

Posted on • Originally published at rentierdigital.xyz

There Are Now Two Internets. You're Only Optimizing for One.

My best article doesn't exist. Well, it does, technically. It's still live, still pulling reads, 25,000 of them and counting, my best performer across the whole catalog. But I just went looking for it on the index an AI agent actually checks when it searches on my behalf, and it wasn't there. Absent. 😬

3 neural search queries on Exa, built straight from the article's own topic. So the question that stays open here is whether that actually matters, and for whom.

My Best Article Isn't There

Here's the setup, in case you want to run it yourself. Exa does neural search, meaning it doesn't match keywords, it matches meaning. I fed it 3 queries pulled straight from my article's own territory, the CLI-versus-MCP debate for AI agents, which is exactly what the piece is about. 10 results per query, 30 results total.

None of them pointed back to me. None of them pointed to Medium as a platform, period.

This isn't some forgotten post buried on page 12 of my catalog. It's my single best performer, the one other articles get measured against, and a CTR most of my posts can only dream of (I run the numbers monthly, this one still tops the list). If an AI agent went looking for exactly what this article covers, it wouldn't find it anywhere in the 10 results per query I pulled.

Checked twice. Didn't believe it the first time either. Second run felt like the world's smallest "YOU DIED" screen.

So what shows up instead of me?

What Shows Up Instead

Not really competitors, not in the way I expected.

The 30 results split into 3 buckets. A pile of GitHub repos, small shim tools with names like "onlycli" and "mcpshim," the kind of project that has 40 stars and a README written last month. A stack of engineering blogs from adjacent platforms, Speakeasy, Courier, Arize, GitHub's own blog, all writing about the exact same tension my article covers. And official documentation, straight from Anthropic's own claude-code-action repo.

On the broader query, a couple of product sites joined the mix (withvibe.dev, amux.io), and Mistral showed up once. Mistral. Not the company I was expecting to compete with on a CLI-versus-MCP query, but there it was.

What's conspicuously missing from all 3 buckets: Medium. Not just my article, the platform itself. I'm not ready to call that structural yet (that comes later), but it's the kind of absence that makes you sit up.

I built the first version of this test after reading the same test on a different set of assistants, which asked a version of this exact question about a different product entirely. Different tool, same blind spot.

Two Engines, Two Realities

None of this is a Medium problem, or an Exa problem. It's a plumbing problem, and once you see the plumbing, the 30 empty results stop being weird and start being expected.

Google indexes by keywords and backlinks. PageRank, in its bones, was built to answer a short query typed by a human into a search bar, ranked by how many other pages point to yours and how well your text matches the words someone typed. Exa, and the wave of competitors chasing the same bet, indexes by semantic similarity on embeddings. It's built to answer a question posed in natural language by an agent that's browsing the web on someone else's behalf, not typing 3 keywords and scanning 10 blue links.

These aren't 2 flavors of the same system. They're 2 disjointed systems built on separate indexes, with different criteria for what counts as a good answer. Pick your pill, they don't talk to each other. And here's the part that actually keeps me up: a piece of content can exist fully in one and be completely absent from the other, and nothing tells the author that happened. No warning email, no dashboard flag, no drop in a metric you're already watching. You'd have to go looking, the exact way I went looking, on a random afternoon, for no better reason than curiosity about my own numbers. If I hadn't run these 3 queries out of idle curiosity, I'd still believe my best article was universally findable, because on the index I check every day (Google Search Console, Medium's own stats), it clearly is. It's just invisible on the other one, and I had zero way of knowing that until I went and asked.

This Isn't a Niche Experiment

3 queries and 30 results is a personal anecdote. The scale behind it isn't.

Exa's CEO, Will Bryk, said on X a few days before I ran this test that the company now serves 80 billion pages and tracks 1.4 trillion URLs, with a stated goal of reaching Google's scale by early 2027. He put numbers on the competition too: Google around 1 trillion pages indexed, Bing around 500 billion, Yandex around 200 billion, Brave around 40 billion. Those are Bryk's own estimates, worth flagging as such, but they're not coming from nowhere. An independent review from The AI Agent Index cross-checked the funding and index-size claims in July and landed on a broadly similar picture, over 500 billion URLs crawled.

<figure data-source="infographic" data-gen-id="nd75nwy8fhsnv789vvz85mvjns8byr31">
TITLE "Whose Index Is Bigger" + subtitle "Search engines built for agents vs search engines built for humans". Metaphor: a set of vertical filing cabinets of different heights, each labeled with a search engine name, standing side by side on a shared shelf. Style: engineer blueprint, blue-line technical drawing on cream paper, thin precise linework, dimension markers. Palette: navy #14213D, amber #FCA311, muted red #C1121F, cream #FFF8E7, charcoal #2B2B2B. Content: five cabinets labeled GOOGLE (tallest, marked approximately 1 trillion pages), BING (approximately 500 billion), YANDEX (approximately 200 billion), EXA (approximately 1.4 trillion tracked URLs, rising fastest, marked with an upward arrow), BRAVE (shortest, approximately 40 billion). Highlight: the EXA cabinet outlined in amber with a dashed extension line above it labeled "TARGET: GOOGLE SCALE, EARLY 2027". Legend: small note bottom-left, "figures self-reported by each company, order of magnitude only". Footer: © rentierdigital.xyz bottom-right, small, handwritten. NOT flat corporate vector, NOT glossy 3D bar chart.\

Search Engine Index Size Comparison: Agents vs Humans

And it's not standing still. In May, Exa closed a $250 million Series C, led by Andreessen Horowitz with Nvidia and Thrive Capital in the round, pushing its valuation to $2.2 billion, tripled in 6 months. Total raised across 4 rounds sits at $361 million. That's not a side project for curious developers anymore. That's infrastructure being built at a sprint, with competitors chasing the same thesis right behind it (Parallel, You.com, and I'd bet money on more showing up before this article's a year old).

If it's already this size and growing this fast, what does that actually change for someone publishing content right now?

This Isn't Just a Developer Problem

I can already hear the objection. This is an API thing, it's a devs-wiring-agents-together thing, it doesn't touch anyone who just writes.

Here's why that's wrong. AI agents are becoming the audience doing the searching, not just a tool bolted onto someone else's product. An assistant answering a question, doing research, recommending a solution, all of it leans on an index like this one, not on Google. Classic SEO, keywords, backlinks, domain authority, optimizes for a system a growing share of searches never touch anymore. Classic "works on my machine" energy, except the machine is Google and the actual deploy target is everything else.

If your content strategy still starts and ends with keyword research and backlink outreach, you're optimizing for 1 internet while the mistakes that keep content out of AI search quietly happen on the other one.

That's where this stops being about my article specifically. It matters, and it matters first for anyone producing content meant to be found, not just people writing code that talks to agents.

Top Google and still be invisible to the internet that's reading for you.

What the Test Doesn't Prove

Now the part I'd rather skip but shouldn't.

3 queries, 30 results, 1 date, 1 topic. That's not an audit. It doesn't prove Medium is structurally locked out of Exa's index in general, only that on these 3 specific subjects, on this specific day, nothing from the platform surfaced. I think that distinction actually matters more than the headline number, though maybe I'm reading too much into a single afternoon of testing.

What I genuinely don't know: is this about crawl freshness, meaning Exa just hasn't gotten around to Medium's newer pages yet? Is it about the topics I picked, maybe CLI-versus-MCP content skews toward GitHub and docs by nature, regardless of platform? Or is there a structural bias baked into the index itself, favoring technical sources over blogging platforms as a category? I don't have the data to pick 1 of those 3 explanations over the others, and pretending I did would be going further than what I actually tested.

Small tangent, unrelated, but it's what I was doing while this test was running in a second tab: I spent 20 minutes trying to figure out why my own analytics dashboard had started showing traffic from a country I don't sell in, and it turned out to be a scraper bot with a residential IP block, not a customer. Not related to any of this. Just what Tuesday afternoon looked like.

Back to the point. What I know: 3 queries, on a subject I know cold, never surfaced my own work. What I don't know: whether that's a platform issue or a topic issue. I'm not going to dress up an afternoon of testing as an audit just because the result is dramatic.

Two Internets, Diverging

2 internets are building themselves in parallel right now, with 2 different sets of rules for what gets discovered, and most people publishing content have no idea which one they're actually filling up.

I'll run this again in a few months, different topics, wider net. For now, that's everything I've got.

Sources

  • Will Bryk (Exa CEO), index-scale claims via X, corroborated independently by The AI Agent Index in July 2026
  • Exa's official Series C announcement, May 20, 2026 (self-published, treated here as company communication, not independent audit)
  • ChatForest, coverage of Exa's $250M Series C and valuation, May 23, 2026

This post may contain affiliate links. If you click them, I might earn a small commission — costs you nothing, and helps me keep shipping quality articles every day for your reading pleasure.

Top comments (0)