DEV Community

Gissur Runarsson for Bersyn

Posted on • Originally published at bersyn.com

ChatGPT and Claude flatly disagree about which developer tools to recommend. I have the receipts.

There is a comfortable assumption behind most "AI visibility" thinking: that the four big AI models broadly agree, so if you are present in one you are roughly present in all. The scan data says the opposite. The same developer tool can be named in five of five Conversations on Claude and zero of five on ChatGPT. Your AI presence is not a property of your company. It is a property of which model your buyer happens to open.

Between 22 May and 4 June 2026 the Bersyn scan engine ran buyer Conversations across four AI Surfaces — ChatGPT, Claude, Perplexity, Gemini — for six developer-infrastructure companies. Five high-intent buyer questions per Surface, twenty Conversations per company. Here is what the disagreement actually looks like.

The receipts, side by side

This is the same five-question test, run against four models, for each company. The numbers are how many of five Conversations named the company on that Surface.

Company Category ChatGPT Claude Perplexity Gemini
Novu Notifications infrastructure 2/5 5/5 5/5 4/5
Infisical Secrets management 0/5 2/5 3/5 0/5
Convex Backend-as-a-service 0/5 2/5 1/5 0/5
Trigger.dev Background jobs / durable workflows 0/5 1/5 2/5 0/5
Windmill Workflow / internal tooling 0/5 2/5 0/5 0/5
Zuplo API gateway 0/5 0/5 2/5 0/5

Now read the Novu row. On Claude and Perplexity it was named in every single Conversation — a clean 5/5, with AI naming Knock as the alternative. On ChatGPT it was named twice, and AI named Courier instead. A founder looking only at the Claude result would believe Novu has best-in-class AI presence. They would be right, on Claude. On ChatGPT the same company is a coin-flip at best, sitting behind Courier.

That is the entire thesis in one row: a 5/5 on one model and a 2/5 on another, for the identical product, on the identical day.

ChatGPT and Gemini agree on one thing: silence

Look down the ChatGPT and Gemini columns. Outside Novu, every company scored 0/5 on both. Convex, Trigger.dev, Windmill, Zuplo, and Infisical were named zero times in twenty combined ChatGPT-and-Gemini Conversations each.

When AI did not name these companies, here is who it recommended instead:

Company Recommended Instead (ChatGPT) Recommended Instead (Gemini)
Convex Firebase Firebase
Trigger.dev Temporal Temporal
Windmill (named no challenger) (named no challenger)
Zuplo Kong Kong
Infisical HashiCorp Vault HashiCorp Vault
Novu Courier Courier

The incumbents — Firebase, Temporal, Kong, HashiCorp Vault — own the ChatGPT and Gemini answer outright in these categories. These two Surfaces lean hardest on training data, and training data rewards the names that have accumulated the most third-party mentions across Reddit, Hacker News, comparison articles, and docs over many years. The newer entrant has not accumulated that mass yet, so on these two Surfaces it does not exist.

Claude and Perplexity tell a different story

The same companies that scored 0/5 on ChatGPT were frequently named on Claude or Perplexity:

  • Convex went from 0/5 on ChatGPT to 2/5 on Claude.
  • Trigger.dev went from 0/5 on ChatGPT to 2/5 on Perplexity.
  • Infisical went from 0/5 on ChatGPT to 3/5 on Perplexity, where it moved out of "Omitted" entirely.
  • Zuplo was named zero times on three Surfaces and twice on Perplexity — its only foothold in the entire scan was the one retrieval-driven model.

Perplexity does live web retrieval at query time, so it surfaces newer tools faster. Claude does some retrieval and is more willing than ChatGPT to name developer-first and open-source tools it has seen in training. That is why the challenger that is invisible on ChatGPT can still be a recommended option on the other two.

The same split appears outside this sample

This is not unique to these six companies. The Bersyn scans from the authentication category over the same window show the identical fault line:

  • SuperTokens scored 8/10 on Claude and 8/10 on Perplexity, and 0/10 on ChatGPT.
  • Hanko was named 5/5 on both Claude and Perplexity, and 0/5 on both ChatGPT and Gemini.
  • Cerbos was named 3/5 on Claude, Perplexity, and Gemini — and 0/5 on ChatGPT, where it named no challenger at all in Cerbos's place.

The pattern repeats across categories: strong on the retrieval-leaning models, invisible on the training-data-leaning ones, for the same company on the same day.

What this means for a founder

You cannot manage what you measure as a single number. "Are we visible in AI?" is the wrong question because it has four different answers. The right questions are:

  1. Which Surface is my buyer most likely to open? (For most B2B buyers today, that is ChatGPT — the harshest Surface in every scan above.)
  2. On that specific Surface, am I named, or is an incumbent named instead?
  3. If I am strong on Claude and Perplexity but invisible on ChatGPT and Gemini, my problem is training-data presence, not retrieval — and the fix is different.

A single average across the four models would have hidden every one of these findings. Novu's average looks healthy; its ChatGPT number is not. Zuplo's average looks hopeless; its Perplexity number is a foothold worth defending.

See where the models disagree about you

If you sell developer infrastructure, the odds are good that at least one of these four models recommends a competitor in your place — and you do not know which one, because you have never seen all four side by side.

Run a free scan at bersyn.com. Two-minute setup, the same scan engine that produced every table above. You will see, model by model, who AI recommended instead of you — and which Surface is quietly handing your category to an incumbent. The disagreement is already happening in front of your buyers. Better to read the receipt than to find out from a lost deal.


Data from Bersyn scans run 22 May – 4 June 2026. Six developer-infrastructure companies: Convex, Trigger.dev, Windmill, Zuplo, Infisical, Novu — plus cross-reference scans for SuperTokens, Hanko, and Cerbos. Five buyer questions per Surface, four Surfaces, twenty Conversations per company. Raw scan JSON available on request. Bersyn does not claim which tool is "best" — the tables above measure which companies the AI Surfaces have learned to recommend, not product quality.

Top comments (0)