<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bersyn</title>
    <description>The latest articles on DEV Community by Bersyn (bersyn).</description>
    <link>https://dev.to/bersyn</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F13431%2F42466cc9-10f3-4036-b770-ae51da8e35ed.png</url>
      <title>DEV Community: Bersyn</title>
      <link>https://dev.to/bersyn</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bersyn"/>
    <language>en</language>
    <item>
      <title>We asked AI for the best CRM 80 times. Attio, the CRM startups love, was named twice.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Fri, 10 Jul 2026 17:45:06 +0000</pubDate>
      <link>https://dev.to/bersyn/we-asked-ai-for-the-best-crm-80-times-attio-the-crm-startups-love-was-named-twice-1k6</link>
      <guid>https://dev.to/bersyn/we-asked-ai-for-the-best-crm-80-times-attio-the-crm-startups-love-was-named-twice-1k6</guid>
      <description>&lt;p&gt;More buyers now ask ChatGPT, Claude, Perplexity, or Gemini for a tool before they ever open Google. So we keep measuring what those models actually say when someone asks for software in a category. CRM turned out to be one of the cleanest examples we have run, because of Attio.&lt;/p&gt;

&lt;p&gt;Attio is the modern CRM that startups reach for and rave about. Well funded, well designed, exactly the tool the online startup crowd talks about. If funding and affection were what the models rewarded, Attio would be named constantly. It is not.&lt;/p&gt;

&lt;p&gt;We asked two versions of the same question and ran each one 10 times on all four models. That is 40 answers per question, 80 in total.&lt;/p&gt;

&lt;p&gt;Version one was the plain category label:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What's the best CRM software for a business?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Version two described the buyer Attio is actually built for, with no tool named:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What should an early-stage startup use to track deals and customer relationships without a heavy enterprise sales tool?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The receipt
&lt;/h2&gt;

&lt;p&gt;How many of the 10 runs named each tool, per model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Best CRM                For an early-stage startup
                    (category label)        (no heavy enterprise tool)
Model               Attio   Salesforce      Attio   Salesforce
-----------------   -----   ----------      -----   ----------
ChatGPT             0/10    10/10           0/10     0/10
Claude              0/10    10/10           2/10     7/10
Gemini              0/10    10/10           0/10     0/10
Perplexity          0/10    10/10           0/10     8/10
-----------------   -----   ----------      -----   ----------
All four models     0/40    40/40           2/40    15/40

Named instead, both questions: HubSpot, Pipedrive, Zoho, and on the
startup question, Notion, Airtable, and Streak.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Attio is not phrasing-sensitive. It is just absent.
&lt;/h2&gt;

&lt;p&gt;This is the part that makes CRM different from other categories we have run. With some tools the naming swings depending on how you ask. Attio does not swing. It was named zero times out of 40 on the plain question and 2 times out of 40 on the startup question. Across all 80 answers it came up twice, both on Claude. The category is owned by Salesforce, HubSpot, Pipedrive, and Zoho, and Attio is simply not in the conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The models can reason about startup fit. They just do not know Attio belongs there.
&lt;/h2&gt;

&lt;p&gt;Here is the tell that this is not the models failing to understand the question. Look at Salesforce on the second question. It went from 40 of 40 down to 15. The models correctly noticed that an early-stage startup that does not want a heavy enterprise tool probably should not be handed Salesforce, so they dropped it and reached for lighter options, HubSpot, Pipedrive, Notion, Airtable, Streak. The reasoning is right there. They know how to pick a leaner CRM for a smaller team. They just have never learned that Attio is one of the answers.&lt;/p&gt;

&lt;p&gt;That is the cleanest version of a corroboration problem. It is not that the model cannot find a startup CRM. It is that the specific sentence, Attio is the modern CRM for an early-stage startup, does not exist widely enough in the places the model reads for it to reach for that name.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means if you are the challenger
&lt;/h2&gt;

&lt;p&gt;Funding does not move the model. Product love does not move the model. A great design that your users adore in private does not become a public sentence the model can learn from. Salesforce and HubSpot are named because two decades of comparison posts, reviews, docs, and threads taught the model they are the answer. Attio has the product. It does not yet have the corpus.&lt;/p&gt;

&lt;p&gt;The work is not a better CRM. It is getting Attio written about, by real users, as the answer to the specific jobs it wins, in the public places these models read. Not more mentions of the category, the exact sentence tied to the job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check your own category
&lt;/h2&gt;

&lt;p&gt;Pick your category. Ask the plain label question and the version that describes the specific job your best customers hire you for. Run both a handful of times on ChatGPT, Claude, Perplexity, and Gemini and count who gets named. If you are the beloved, well-funded challenger who still does not show up, the fix is not the product.&lt;/p&gt;

&lt;p&gt;If you want the fast version, run your own domain through the free scan and see who AI recommends in your category and who it names instead: &lt;a href="https://bersyn.com/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=crm-attio-teardown" rel="noopener noreferrer"&gt;bersyn.com free scan&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In your category, is the tool that everyone online talks about the same one AI actually names when a buyer asks?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>saas</category>
      <category>startup</category>
      <category>marketing</category>
    </item>
    <item>
      <title>We asked AI for the best email tool 40 times. Beehiiv barely showed up. Then we asked about newsletters.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Fri, 10 Jul 2026 17:45:05 +0000</pubDate>
      <link>https://dev.to/bersyn/we-asked-ai-for-the-best-email-tool-40-times-beehiiv-barely-showed-up-then-we-asked-about-44hk</link>
      <guid>https://dev.to/bersyn/we-asked-ai-for-the-best-email-tool-40-times-beehiiv-barely-showed-up-then-we-asked-about-44hk</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Correction, 2026-07-16.&lt;/strong&gt; The original version of this post said the lesson was about framing: reword your category question as the specific job and the challenger wins. A reader, Paul Towers, caught a confound. The two questions below did not only change from a category label to a job. They also changed the buyer, from a business to a creator. So I reran it holding the buyer constant, 10 times per model, and the framing effect disappeared. A business buyer who describes the job in plain words still gets Beehiiv 0 out of 40, the same as the category label. The corrected finding is at the top. The original numbers are left untouched below, for the record.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What actually moved the result: who is asking, not how they phrase it
&lt;/h2&gt;

&lt;p&gt;I reran the test with a third question, keeping the buyer a business the whole time and only changing the framing from a category label to the specific job.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Question                                                    Beehiiv (all 4 models)
--------------------------------------------------------    ----------------------
Business buyer, category label                              3/40
  "What's the best email marketing platform for a business?"
Business buyer, the job in plain words                      0/40
  "We have an existing customer base we want to reach via
   an owned channel like email. Walk me through it."
Creator buyer, the job                                      25/39
  "How does a creator or founder start a newsletter, grow
   subscribers, and make money from it?"
--------------------------------------------------------    ----------------------
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Holding the buyer constant and changing only the framing takes Beehiiv from 3 to zero. There is no framing effect. The entire jump was the buyer identity. Beehiiv is filed under "creator," and it vanishes for a business buyer at any phrasing. ChatGPT named it 0 out of 10 on all three questions, which is the one part of the original that survives.&lt;/p&gt;

&lt;p&gt;The real lesson is harder and truer than "phrase it as a job." You cannot reframe your way into a category. The models file you under a buyer identity, written in the words the people who actually use you have published about you. Beehiiv owns the creator's newsletter job because the creator world documented it as the answer to that job. A business buyer never reaches it, because nobody wrote Beehiiv up as the answer to a business's email program, no matter how that business phrases the question.&lt;/p&gt;

&lt;p&gt;So if you are the challenger: the move is not to reword your question. It is to pick the specific buyer whose job you can genuinely own, and get documented, by real users and real publications, as the answer to that job in the words that buyer uses. Then the models reach for you when that buyer asks, and only when that buyer asks. Watch ChatGPT separately, because it moves last.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Original post below, kept for the record.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;More buyers now ask ChatGPT, Claude, Perplexity, or Gemini for a tool before they ever open Google. So we keep measuring what those models actually say when someone asks for software in a category. Email marketing gave us the most encouraging result we have run, because of Beehiiv. It is the one teardown where the challenger actually won its slot, and you can see exactly how.&lt;/p&gt;

&lt;p&gt;We asked two versions of the same question and ran each one 10 times on all four models. That is 40 answers per question.&lt;/p&gt;

&lt;p&gt;Version one was the plain category label:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What's the best email marketing platform for a business?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Version two was the specific job Beehiiv is built for:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How does a creator or founder start a newsletter, grow subscribers, and make money from it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The receipt
&lt;/h2&gt;

&lt;p&gt;How many of the 10 runs named each tool, per model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Best email tool         Start and grow a newsletter
                    (category label)        (the job Beehiiv is built for)
Model               Beehiiv  Mailchimp      Beehiiv  Mailchimp
-----------------   -------  ---------      -------  ---------
ChatGPT             0/10     10/10          0/10     10/10
Claude              0/10     10/10          10/10    10/10
Gemini              0/10     10/10          9/10     10/10
Perplexity          4/10      9/10          10/10     7/10
-----------------   -------  ---------      -------  ---------
All four models     4/40     39/40          29/40    37/40

Named on the generic question: Mailchimp, Klaviyo, ActiveCampaign, Brevo,
Constant Contact. Named on the newsletter question: Beehiiv, Kit, Substack.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  On the category label, Beehiiv is almost invisible
&lt;/h2&gt;

&lt;p&gt;Ask for the best email marketing platform and Beehiiv was named 4 times out of 40, all of them on Perplexity. Mailchimp was named 39 times, with Klaviyo, ActiveCampaign, Brevo, and Constant Contact filling the rest. If Beehiiv tried to win the generic email marketing question, it would lose, the same way every challenger loses the category label.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one model that still will not budge
&lt;/h2&gt;

&lt;p&gt;Look at ChatGPT. Even on the newsletter question, the question written for Beehiiv's exact buyer, ChatGPT named it 0 times out of 10. This keeps happening in our teardowns. ChatGPT is consistently the slowest of the four to name a challenger, even one that the other three now reach for easily. If ChatGPT is where your buyers ask, winning the sub-intent on the other three is not enough on its own, and it is worth knowing that before you assume the job is done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check your own category
&lt;/h2&gt;

&lt;p&gt;Pick your category. Ask the plain label question, then ask it the way your actual customers describe the job they hire you for, in their own words. Run both a handful of times on ChatGPT, Claude, Perplexity, and Gemini and count who gets named. If you are invisible on the label but present when the question sounds like your real buyer, that tells you which buyer the models have you filed under. If you are invisible on both, &lt;br&gt;
Gemini              0/10     10/10          9/10     10/10&lt;br&gt;
Perplexity          4/10      9/10          10/10     7/10&lt;/p&gt;




&lt;p&gt;All four models     4/40     39/40          29/40    37/40&lt;/p&gt;

&lt;p&gt;Named on the generic question: Mailchimp, Klaviyo, ActiveCampaign, Brevo,&lt;br&gt;
Constant Contact. Named on the newsletter question: Beehiiv, Kit, Substack.&lt;/p&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;


## On the category label, Beehiiv is almost invisible

Ask for the best email marketing platform and Beehiiv was named 4 times out of 40, all of them on Perplexity. Mailchimp was named 39 times, with Klaviyo, ActiveCampaign, Brevo, and Constant Contact filling the rest. If Beehiiv tried to win the generic email marketing question, it would lose, the same way every challenger loses the category label.

## The one model that still will not budge

Look at ChatGPT. Even on the newsletter question, the question written for Beehiiv's exact buyer, ChatGPT named it 0 times out of 10. This keeps happening in our teardowns. ChatGPT is consistently the slowest of the four to name a challenger, even one that the other three now reach for easily. If ChatGPT is where your buyers ask, winning the sub-intent on the other three is not enough on its own, and it is worth knowing that before you assume the job is done.

## Check your own category

Pick your category. Ask the plain label question, then ask it the way your actual customers describe the job they hire you for, in their own words. Run both a handful of times on ChatGPT, Claude, Perplexity, and Gemini and count who gets named. If you are invisible on the label but present when the question sounds like your real buyer, that tells you which buyer the models have you filed under. If you are invisible on both, you have a job to go own.

If you want the fast version, run your own domain through the free scan and see who AI recommends in your category: [bersyn.com free scan](https://bersyn.com/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=email-beehiiv-teardown).

Which buyer are the models filing you under, and is it the one you actually sell to?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>saas</category>
      <category>marketing</category>
      <category>startup</category>
    </item>
    <item>
      <title>We asked AI for the best project management tool 40 times. Linear, the one engineers love, was named twice.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Fri, 10 Jul 2026 17:44:04 +0000</pubDate>
      <link>https://dev.to/bersyn/we-asked-ai-for-the-best-project-management-tool-40-times-linear-the-one-engineers-love-was-3k4m</link>
      <guid>https://dev.to/bersyn/we-asked-ai-for-the-best-project-management-tool-40-times-linear-the-one-engineers-love-was-3k4m</guid>
      <description>&lt;p&gt;More buyers now ask ChatGPT, Claude, Perplexity, or Gemini for a tool before they ever open Google. So we keep measuring what those models actually say when someone asks for software in a category. Project management turned out to be one of the sharpest examples we have run, because of one tool: Linear.&lt;/p&gt;

&lt;p&gt;Linear is a darling among engineers. Fast, opinionated, built for exactly the teams that talk about their tools online. If adoption and affection were what the models rewarded, Linear would be named constantly. It is not.&lt;/p&gt;

&lt;p&gt;We asked two versions of the same question and ran each one 10 times on all four models, so a single lucky or unlucky answer could not fool us. That is 40 answers per question.&lt;/p&gt;

&lt;p&gt;Version one was the plain category label:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What's the best project management software for a team?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Version two was the actual job, phrased the way the buyer Linear is built for would phrase it, with no tool named:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How does a fast-moving engineering team track issues and plan sprints with as little process overhead as possible?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The receipt
&lt;/h2&gt;

&lt;p&gt;How many of the 10 runs named Linear, and named Jira, on each model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Best PM software        Fast engineering team
                    (category label)        (the actual job)
Model               Linear    Jira          Linear    Jira
-----------------   ------    ----          ------    ----
ChatGPT             0 / 10    10 / 10        0 / 10    10 / 10
Claude              2 / 10    10 / 10        5 / 10     9 / 10
Gemini              0 / 10    10 / 10        9 / 10    10 / 10
Perplexity          0 / 10    10 / 10        1 / 10    10 / 10
-----------------   ------    ----          ------    ----
All four models     2 / 40    40 / 40       15 / 40   39 / 40
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  On the category label, Linear is basically invisible
&lt;/h2&gt;

&lt;p&gt;Ask for the best project management software and Linear was named 2 times out of 40. On ChatGPT, Gemini, and Perplexity it was a flat zero. Jira was named on all 40 runs, and behind it the models handed back the same incumbents every time: Asana, Monday.com, Trello, ClickUp, Notion. The category label is settled, and Linear is not in the settlement, no matter how many engineering teams love it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reframe to Linear's actual job and it appears, but only on some models
&lt;/h2&gt;

&lt;p&gt;When we asked the second question, the one that describes what Linear is actually for, Linear climbed to 15 of 40. But look at where the movement came from. Gemini went from 0 to 9. Claude went from 2 to 5. Perplexity barely moved, 0 to 1. And ChatGPT stayed at a flat zero even on the question written for Linear's exact buyer.&lt;/p&gt;

&lt;p&gt;So the tool that engineers reach for first is a tool the biggest model will not name, even when you describe its job precisely. Meanwhile Jira held at 39 and 40. The incumbent survives both questions. The challenger only exists on the phrasing that matches its job, and only on the models willing to name it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means if you are the challenger
&lt;/h2&gt;

&lt;p&gt;Two things worth taking from this.&lt;/p&gt;

&lt;p&gt;The category label is not your fight. If you sell into a category with an entrenched incumbent, the generic question is already lost, and asking it harder will not change that. Linear is proof: real adoption, real affection, still 2 of 40 on the plain question.&lt;/p&gt;

&lt;p&gt;Adoption does not automatically become recommendation. A model names what has been written down in a form it can read, in the words the buyer uses. Engineers picking Linear in a Slack thread does not teach the model anything. The opening is the job-framed question, the specific sub-intent where you actually fit, published clearly enough and corroborated widely enough that a model reaches for you there. That is the sentence worth owning, not the category label you will lose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check your own category
&lt;/h2&gt;

&lt;p&gt;Pick your category. Write the plain label version and the version that describes the specific job your best customers hire you for. Run both a handful of times on ChatGPT, Claude, Perplexity, and Gemini and watch who gets named as the wording changes. If you are the beloved challenger who never shows up on the generic question, you are in good company, and the fix is not a better product.&lt;/p&gt;

&lt;p&gt;If you want the fast version, run your own domain through the free scan and see who AI recommends in your category and how the wording moves it: &lt;a href="https://bersyn.com/?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=pm-linear-teardown" rel="noopener noreferrer"&gt;bersyn.com free scan&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In your category, does the tool everyone actually uses match the tool AI actually names, or are those two different lists?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>saas</category>
      <category>productivity</category>
      <category>startup</category>
    </item>
    <item>
      <title>We asked AI to recommend tools in three SaaS categories 20 times each, and the shape of the answers told us which ones a challenger can still win</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Thu, 09 Jul 2026 16:30:11 +0000</pubDate>
      <link>https://dev.to/bersyn/we-asked-ai-to-recommend-tools-in-three-saas-categories-20-times-each-and-the-shape-of-the-answers-2fec</link>
      <guid>https://dev.to/bersyn/we-asked-ai-to-recommend-tools-in-three-saas-categories-20-times-each-and-the-shape-of-the-answers-2fec</guid>
      <description>&lt;p&gt;More buyers now ask ChatGPT, Claude, Perplexity, or Gemini for a tool before they ever open Google. So we started measuring what those models actually say when someone asks for software in a category. This week we ran a simple probe across three developer categories, and the result was less about who won and more about the shape of each race. Some categories are frozen at the top. Some are wide open. And the shape tells a founder what is even worth attempting.&lt;/p&gt;

&lt;p&gt;We took three categories, asked one plain buyer question in each, and ran that question 5 times on each of the four models. That is 20 answers per category. We counted how many of the 20 named each company.&lt;/p&gt;

&lt;h2&gt;
  
  
  The receipt
&lt;/h2&gt;

&lt;p&gt;Here are the three categories, the top named tools, and how many of the 20 answers named each one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;20 answers per category (5 runs x 4 models: ChatGPT, Claude, Perplexity, Gemini)

Product analytics          Feature flags              Error monitoring
(frozen at the top)        (wide open)                (owned, but deep tail)
-----------------------    -----------------------    -----------------------
Mixpanel      20 / 20      LaunchDarkly  20 / 20      Sentry        20 / 20
Amplitude     20 / 20      Flagsmith     17 / 20      Bugsnag       19 / 20
Heap          19 / 20      Optimizely    17 / 20      Rollbar       15 / 20
PostHog       15 / 20      Split         15 / 20      Raygun        14 / 20
--- cliff ---              GrowthBook    13 / 20      LogRocket     11 / 20
Segment        7 / 20      Unleash       13 / 20      Datadog        6 / 20
Hotjar         6 / 20      Statsig       11 / 20
FullStory      5 / 20      PostHog       11 / 20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the three columns as three different shapes, because that is the whole point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Product analytics is frozen
&lt;/h2&gt;

&lt;p&gt;Mixpanel and Amplitude were named on every single answer, all four models, all five runs. Heap held at 19 of 20. PostHog came in at 15, and then there is a cliff: Segment 7, Hotjar 6, FullStory 5. The top of this category is concrete. If you are a new analytics tool, the generic "best product analytics tool" question is not a fight you win right now. The models have already settled on the top two, and asking the plain question harder will not change that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Feature flags is wide open
&lt;/h2&gt;

&lt;p&gt;Now look at the middle column. LaunchDarkly led at 20 of 20, the way you would expect from the incumbent. But under it the field stays alive: Flagsmith 17, Optimizely 17, Split 15, GrowthBook 13, Unleash 13, Statsig 11, PostHog 11. That is eight tools with real, repeated share in the same 20 answers. There is no cliff here. A challenger in this category genuinely gets named, because the models have not collapsed the answer down to two names. If this is your category, the move is not clever. It is to describe what you do clearly enough that a model can place you, because the slot is actually available.&lt;/p&gt;

&lt;h2&gt;
  
  
  Error monitoring is owned, but the tail runs deep
&lt;/h2&gt;

&lt;p&gt;Sentry owned this one at 20 of 20. Bugsnag was right behind at 19, then Rollbar 15, Raygun 14, LogRocket 11. So the top is locked, but unlike product analytics there is no cliff after the leader. Five and six tools deep still get named.&lt;/p&gt;

&lt;p&gt;The most useful number in the whole probe is Datadog at 6 of 20 here. Datadog is everywhere in monitoring. Ask the generic question and it shows up constantly. It came in low in this probe because we asked the question the way a small team asks it, "for a small dev team", and that sub-intent quietly pushed the heavyweight down and let the leaner tools surface. The lock is real on the generic question. It breaks on the sub-intent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the shape tells you to do
&lt;/h2&gt;

&lt;p&gt;The shape of your category tells you what to even attempt.&lt;/p&gt;

&lt;p&gt;If you are in a frozen top-two category like product analytics, do not spend your energy trying to get named next to the incumbents on the generic question. You will not, and the probe shows why: the top two are named on every answer and there is a cliff behind them. The opening is in the sub-intents. The questions that carry a qualifier, "for small teams", "cheapest", "for [specific use case]", are where the lock cracks, the same way Datadog drops out the moment the question says "small dev team". A frozen category is not unwinnable. It is winnable one specific buyer question at a time, not on the category label.&lt;/p&gt;

&lt;p&gt;If you are in an open category like feature flags, the game is different and easier. The models are still deciding, so you get named by being legible. Describe what you actually do, in the words a buyer uses, and a model can place you next to the eight names already in the mix. You do not need a trick. You need to be clear while the category is still open.&lt;/p&gt;

&lt;p&gt;Either way, the first thing to know is which shape you are in, because it decides whether you chase the generic question or go hunting for the sub-intent where the incumbents thin out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check your own category
&lt;/h2&gt;

&lt;p&gt;Pick your category. Ask one plain buyer question, run it a handful of times on ChatGPT, Claude, Perplexity, and Gemini, and count who gets named. If the same two names show up on every answer with a cliff behind them, you are frozen, and your opening is in the sub-intents. If a real field of tools keeps appearing, you are open, and your job is to describe yourself clearly enough to join it.&lt;/p&gt;

&lt;p&gt;If you want the fast version, run your own domain through the free scan and see who AI recommends in your category and how deep the field goes: &lt;a href="https://bersyn.com/?utm_source=blog&amp;amp;utm_medium=referral&amp;amp;utm_campaign=frozen-vs-open" rel="noopener noreferrer"&gt;run the free scan&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In your category, is the top frozen at two names, or is the field still open?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>saas</category>
      <category>startup</category>
      <category>marketing</category>
    </item>
    <item>
      <title>We reworded one buyer question and watched a shared inbox tool drop from 10 out of 10 to zero on ChatGPT</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Thu, 09 Jul 2026 15:42:03 +0000</pubDate>
      <link>https://dev.to/bersyn/we-reworded-one-buyer-question-and-watched-a-shared-inbox-tool-drop-from-10-out-of-10-to-zero-on-17f5</link>
      <guid>https://dev.to/bersyn/we-reworded-one-buyer-question-and-watched-a-shared-inbox-tool-drop-from-10-out-of-10-to-zero-on-17f5</guid>
      <description>&lt;p&gt;More and more buyers ask an AI model before they ever open Google. So we started measuring what ChatGPT, Claude, Perplexity, and Gemini actually say when someone asks for a tool in a category. Shared inbox software gave us the cleanest example we have found so far of something that should worry anyone who sells software.&lt;/p&gt;

&lt;p&gt;We asked two versions of the same question. Then we ran each version 10 times on all four models, so a single lucky or unlucky answer could not fool us.&lt;/p&gt;

&lt;p&gt;Version one was the plain category label:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What's the best shared inbox software for small teams?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Version two was how an actual buyer types it when they have the real problem in their head:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What's the best way to collaborate with my team on shared email inboxes without losing our email workflow?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same intent. Same shortlist of tools that could answer it. The only thing that changed was the wording.&lt;/p&gt;

&lt;h2&gt;
  
  
  The receipt
&lt;/h2&gt;

&lt;p&gt;Here is ChatGPT, the same question asked both ways, 10 runs each. The number is how many of the 10 runs named that company.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ChatGPT, 10 runs per question

Company        Category label      Reworded as a buyer
------------   ----------------    -------------------
Missive        10 / 10             0 / 10
Hiver          10 / 10             9 / 10
Front          10 / 10             10 / 10
Help Scout     10 / 10             10 / 10
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the middle column again. On the category label, ChatGPT recommended Missive on every single run. Reword the exact same question the way a real buyer would say it, and Missive was named zero times out of ten. Same company, same website, same 10 runs, one different sentence.&lt;/p&gt;

&lt;p&gt;Hiver barely moved (10 to 9). Front and Help Scout did not move at all on ChatGPT. They held a perfect 10 out of 10 through the rewording.&lt;/p&gt;

&lt;h2&gt;
  
  
  It was not just ChatGPT, and not just Missive
&lt;/h2&gt;

&lt;p&gt;Perplexity produced the mirror image. On the category label it named Hiver on all 10 runs. On the reworded question it named Hiver zero times, and started pointing people toward doing it in Gmail and Outlook instead of naming a product at all.&lt;/p&gt;

&lt;p&gt;The direction of the swing depended on the model. On Claude the reworded question pulled Hiver down (8 to 5) and nudged Missive down (10 to 8). On Gemini the reworded question actually lifted Hiver (6 to 9) while Missive held (9 to 8). Front and Help Scout were the steadiest of the four, holding 10 out of 10 across ChatGPT, Claude and Gemini, and only softening on Perplexity's reworded question, where almost everything dropped because Perplexity stopped naming products. There was no single pattern across models. The one thing that stayed calm everywhere was the plain category question. The reworded question is where companies appeared and disappeared.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things this should change about how you check
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A single scan will lie to you.&lt;/strong&gt; If we had asked ChatGPT once, seen Missive missing, and stopped there, we would have walked away certain Missive had some deep AI problem. They do not. They are recommended on 10 of 10 runs on the other phrasing. One question, on one model, on one run, tells you close to nothing. The behavior only shows up across runs and across phrasings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The companies that held did not do anything clever.&lt;/strong&gt; Front and Help Scout were not gaming a model. Their pages describe the actual job a buyer is trying to do, in the words a buyer uses, so a model can place them no matter how the question is asked. The companies that only carry the category label survive the label question and come apart on the workflow one. This is a legibility problem, not a trick. It is whether your site says what you do in the buyer's own words clearly enough that a model can recommend you across every way the question gets asked.&lt;/p&gt;

&lt;p&gt;If you sell software, your category label is the question you already rank for in your head. The reworded, real-buyer version is the one you have probably never checked. That is the one deciding whether AI names you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check your own category
&lt;/h2&gt;

&lt;p&gt;Pick your category. Write the plain label version and the way one of your actual customers would ask it. Run both a handful of times on ChatGPT, Claude, Perplexity, and Gemini, and watch which of your competitors appear and disappear as the wording changes.&lt;/p&gt;

&lt;p&gt;If you want the fast version, you can run your own domain through the free scan and see who AI recommends in your category and how the wording moves it: &lt;a href="https://bersyn.com/?utm_source=blog&amp;amp;utm_medium=referral&amp;amp;utm_campaign=shared-inbox-teardown" rel="noopener noreferrer"&gt;run the free scan&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In your category, is the recommendation this sensitive to phrasing, or did shared inbox software just happen to be an unusually swingy one?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>saas</category>
      <category>marketing</category>
      <category>seo</category>
    </item>
    <item>
      <title>Why AI recommends the same tools every time, and which slots you can actually win</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Tue, 30 Jun 2026 12:52:11 +0000</pubDate>
      <link>https://dev.to/bersyn/why-ai-recommends-the-same-tools-every-time-and-which-slots-you-can-actually-win-850</link>
      <guid>https://dev.to/bersyn/why-ai-recommends-the-same-tools-every-time-and-which-slots-you-can-actually-win-850</guid>
      <description>&lt;p&gt;Last week I &lt;a href="https://www.bersyn.com/blog/ai-recommends-neon-for-databases-specialists-invisible-2026" rel="noopener noreferrer"&gt;tore down what four AI models recommend for databases&lt;/a&gt;: same buyer questions, 20 runs each, count who gets named. Neon, Upstash and Turso already own the generic slots. The specialists, Tigris and Tinybird, are close to invisible.&lt;/p&gt;

&lt;p&gt;The post did fine. The comments did better. A handful of sharp people pushed on the &lt;em&gt;why&lt;/em&gt;, and the answer is more useful than the findings were.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two competitions, and you are probably in the wrong one
&lt;/h2&gt;

&lt;p&gt;The models are not judging which product is best. They surface whatever their training data already corroborated for that exact phrasing. A worse tool with denser, clearer writing tied to a question beats a better tool nobody wrote about that way.&lt;/p&gt;

&lt;p&gt;So there are two competitions running at once:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best product.&lt;/strong&gt; What you actually build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best-corroborated answer.&lt;/strong&gt; What the sources a buyer's question pulls from already named.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most founders pour everything into the first and assume the second follows. It does not. The model only scores the second.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some slots are welded shut. Stop fighting them.
&lt;/h2&gt;

&lt;p&gt;"Object storage" has meant Amazon S3 for fifteen years. That mental model is set in concrete across millions of pages. No comparison article, no migration guide, no amount of content moves "best object storage" off S3 on any timeline that matters to a startup. If your growth plan depends on winning a query like that, the plan is the problem.&lt;/p&gt;

&lt;p&gt;The tell for a welded-shut slot: ask the four models the same generic question a few times and they all agree, every run. Agreement across models and across runs means consensus has formed. You are not getting in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some slots are wide open. That is where the work pays.
&lt;/h2&gt;

&lt;p&gt;Now ask a narrow sub-job instead. "Vector database for X." "Analytics for Y." Watch what happens: the models hedge, name different tools, and the first pick shuffles run to run. That disagreement is the signal. Consensus has not formed yet, which means the slot is still being decided, which means you can be the one it decides on.&lt;/p&gt;

&lt;p&gt;The challengers that won did exactly this. Neon, Upstash and Turso did not beat Postgres at "best database." They became the corroborated answer for "serverless Postgres / Redis / SQLite" while those mental models were still forming, and rode them as they widened.&lt;/p&gt;

&lt;p&gt;So the move in a locked category is not to attack the incumbent's query. It is to find the sub-job nobody owns, become the best-corroborated answer for it in the buyer's own words, and let the category grow around you. Manufacture a young category you can actually win.&lt;/p&gt;

&lt;h2&gt;
  
  
  You can see which is which
&lt;/h2&gt;

&lt;p&gt;This is the part founders miss: you do not have to guess whether a slot is welded or winnable. You can observe it. Run the buyer questions across the models, more than once, and look at the agreement. Tight agreement across models and runs is a settled slot. Disagreement is an open one. That turns "get recommended by AI" from a vibe into a map of where to spend.&lt;/p&gt;

&lt;p&gt;That map is what &lt;a href="https://www.bersyn.com" rel="noopener noreferrer"&gt;Bersyn&lt;/a&gt; builds: who each model names, who gets recommended first, where they disagree, and the verbatim answers behind every number. If you want it for your own category, that is the whole product.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Method: category-representative buyer questions across ChatGPT, Claude, Gemini and Perplexity, multiple runs, reported with model versions and scan dates. No claim that any tool is good or bad, only what the models answered.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>marketing</category>
      <category>startup</category>
      <category>seo</category>
    </item>
    <item>
      <title>I asked four AI models which database to use in 2026. Neon already won. Four challengers are invisible.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Mon, 29 Jun 2026 07:40:24 +0000</pubDate>
      <link>https://dev.to/bersyn/i-asked-four-ai-models-which-database-to-use-in-2026-neon-already-won-four-challengers-are-533g</link>
      <guid>https://dev.to/bersyn/i-asked-four-ai-models-which-database-to-use-in-2026-neon-already-won-four-challengers-are-533g</guid>
      <description>&lt;p&gt;Every week I take one buyer category, ask ChatGPT, Claude, Gemini and Perplexity the five questions a real buyer would type, and count who gets named and who gets recommended first. Same questions for every company, so it is a fair board and not a vibe.&lt;/p&gt;

&lt;p&gt;This week: databases and storage. I expected the incumbent reflex. I got the opposite, then a twist.&lt;/p&gt;

&lt;h2&gt;
  
  
  The serverless newcomers already won
&lt;/h2&gt;

&lt;p&gt;To the models, the challengers are already the answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Neon&lt;/strong&gt; (serverless Postgres): recommended first in 14 of 20 conversations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upstash&lt;/strong&gt; (serverless Redis): first in 11 of 20.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turso&lt;/strong&gt; (edge SQLite): first in 9 of 20.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is real proof the door is not locked. A company younger than the incumbent it replaced can become AI's default pick. Neon did it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I asked about the specialized jobs
&lt;/h2&gt;

&lt;p&gt;Same category, different sub-job, and the model snaps back to the incumbent every time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Vector search:&lt;/strong&gt; the models pick Pinecone, Milvus and Weaviate. &lt;a href="https://www.bersyn.com/recommends/databases-storage/qdrant" rel="noopener noreferrer"&gt;Qdrant&lt;/a&gt; is named a lot but recommended first only six times.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Object storage:&lt;/strong&gt; the answer is Amazon S3, Cloudflare R2 and Backblaze. &lt;a href="https://www.bersyn.com/recommends/databases-storage/tigris" rel="noopener noreferrer"&gt;Tigris&lt;/a&gt; is invisible, named zero times on ChatGPT, Claude and Gemini.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time analytics:&lt;/strong&gt; ClickHouse, not &lt;a href="https://www.bersyn.com/recommends/databases-storage/tinybird" rel="noopener noreferrer"&gt;Tinybird&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The ORM:&lt;/strong&gt; Prisma, not &lt;a href="https://www.bersyn.com/recommends/databases-storage/drizzle" rel="noopener noreferrer"&gt;Drizzle&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why
&lt;/h2&gt;

&lt;p&gt;The newcomers that won did not win on features. By the time these models trained, enough independent writing already named them as the answer. They are in the inputs the model reads.&lt;/p&gt;

&lt;p&gt;The invisible ones shipped great products, and it does not register, because the model is not evaluating products. It repeats what its inputs said. You cannot test your way into a recommendation you were never part of, and you cannot out-feature your way in either. The buyer who types "best vector database" and takes the first answer never sees Qdrant, no matter how good it is, until the inputs change.&lt;/p&gt;

&lt;h2&gt;
  
  
  See it yourself
&lt;/h2&gt;

&lt;p&gt;The full board, with every model's answer and the verbatim text, is here: &lt;a href="https://www.bersyn.com/recommends/databases-storage" rel="noopener noreferrer"&gt;What AI recommends for Databases and Storage&lt;/a&gt;. Counts only, no score, every number links to the actual answer.&lt;/p&gt;

&lt;p&gt;If you want the same teardown for your own category, that is what &lt;a href="https://www.bersyn.com" rel="noopener noreferrer"&gt;Bersyn&lt;/a&gt; does.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>database</category>
      <category>webdev</category>
      <category>saas</category>
    </item>
    <item>
      <title>I asked four AI models which observability tool to use in 2026. They keep naming Datadog and Splunk, and never Better Stack.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Fri, 26 Jun 2026 11:19:40 +0000</pubDate>
      <link>https://dev.to/bersyn/i-asked-four-ai-models-which-observability-tool-to-use-in-2026-they-keep-naming-datadog-and-5fkg</link>
      <guid>https://dev.to/bersyn/i-asked-four-ai-models-which-observability-tool-to-use-in-2026-they-keep-naming-datadog-and-5fkg</guid>
      <description>&lt;p&gt;We build Bersyn, a tool that tracks which products AI models name when someone asks for a recommendation in a category. So we run a lot of scans. This is the third category in a row where the same thing happened, and observability is the cleanest example yet, so it is worth showing the receipts.&lt;/p&gt;

&lt;p&gt;The question is the one a real engineer types into ChatGPT: "what is the best platform to monitor and debug my B2B SaaS app in production, and what are the strong alternatives?" We asked it five ways across four Surfaces: ChatGPT, Claude, Gemini and Perplexity. Then we measured how often each modern tool actually got named, each one scanned in its own home category.&lt;/p&gt;

&lt;h2&gt;
  
  
  The modern tools are not in the answer
&lt;/h2&gt;

&lt;p&gt;Here is the Recommendation Share for five tools that engineers actually talk about, measured across the four Surfaces.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                ChatGPT   Claude   Perplexity   Gemini
Better Stack       0%       0%         0%          0%
Axiom              0%      40%         0%          0%
Highlight          0%      40%         0%          0%
OpenStatus         0%       0%        80%          0%
Checkly           20%      60%        40%         20%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the Better Stack row again. Zero, zero, zero, zero. Not a low score, an absence on every Surface we tested. Axiom and Highlight are named by exactly one model, Claude, and by none of the other three. OpenStatus exists only on Perplexity. A buyer who opens ChatGPT, which is most of them, walks away from this category never having heard four of these five names.&lt;/p&gt;

&lt;h2&gt;
  
  
  What AI names instead
&lt;/h2&gt;

&lt;p&gt;So who got recommended in their place? Here is the tool AI reached for first when each modern challenger was not named.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Better Stack   -&amp;gt;  Splunk, Datadog, Loggly
Axiom          -&amp;gt;  Datadog, Honeycomb
Highlight      -&amp;gt;  FullStory, LogRocket, Sentry
OpenStatus     -&amp;gt;  Cachet
Checkly        -&amp;gt;  Pingdom, Datadog Synthetics, Grafana
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here is ChatGPT, verbatim, asked for the best log management and uptime monitoring platform, the exact category Better Stack sells into:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Choosing the best log management and uptime monitoring platform for a B2B SaaS team depends on several factors... ### Log Management Platforms 1. &lt;strong&gt;Splunk&lt;/strong&gt; - Pros: Highly scalable, powerful search capabilities, extensive integrations, and strong data visualization tools.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And here is ChatGPT on observability and log management, the category Axiom sells into:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Here are some of the top platforms that are widely recognized for their capabilities in observability and log management: 1. &lt;strong&gt;Datadog&lt;/strong&gt;: Datadog is a comprehensive monitoring and analytics platform for developers, IT operations teams, and businesses...&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Splunk. Datadog. Pingdom. Cachet. LogRocket. Look at that list. These are the names that dominated the monitoring conversation around 2015 to 2018, when the training data was thick. Each modern challenger loses, in its own home category, to an incumbent from the era before it existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is not AI being clueless about the category
&lt;/h2&gt;

&lt;p&gt;Here is the part that makes it a real problem rather than a funny one. AI is not ignorant of observability. Ask it and it confidently knows two newer names:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                ChatGPT   Claude   Perplexity   Gemini
Sentry           100%      80%        60%         60%
Honeycomb         40%      40%        80%         40%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sentry gets named in every single ChatGPT answer. Honeycomb shows up across all four models. So the models have room for modern tools in their mental map of monitoring. They have simply frozen that map around the incumbents plus the one or two challengers that broke through years ago. Everything that arrived after the map froze is Omitted.&lt;/p&gt;

&lt;p&gt;That gap has a name in our world: Model Disagreement. When Axiom is named by Claude and by none of the other three, it is not "low visibility." It is invisible on the Surfaces most buyers use, and visible on the one they use least.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this should bother a founder, not just amuse them
&lt;/h2&gt;

&lt;p&gt;The reflex is to wave it off. AI is behind, the models have a training cutoff, it will catch up. Maybe. But your buyer is asking the question today, and the answer they get today hands them Datadog.&lt;/p&gt;

&lt;p&gt;Search had twenty years to learn that Better Stack exists. AI recommendation answers are being formed right now, off whatever evidence the models can find, and for newer companies that evidence is thin. So the incumbent gets named by reflex and the challenger gets skipped.&lt;/p&gt;

&lt;p&gt;The four ways a company shows up wrong in these answers, in our vocabulary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Omitted. The model lists competitors and skips you. This is Better Stack, Axiom and Highlight on ChatGPT.&lt;/li&gt;
&lt;li&gt;Misclassified. The model files you under the wrong category.&lt;/li&gt;
&lt;li&gt;Generic. The model names you so vaguely no buyer could shortlist you.&lt;/li&gt;
&lt;li&gt;Confused. The model conflates you with a similarly named competitor.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What actually moves it
&lt;/h2&gt;

&lt;p&gt;We do not claim AI is biased or that anyone is paying for placement. We read what the models say and show you the evidence. What changes the answer over time is the same thing that changed search: published, specific, verifiable evidence that associates your company with the category and the buyer question. Comparison pages that name the alternatives honestly. Documentation that states plainly what you are and who you are for. Third party mentions in the exact words an engineer would use when asking.&lt;/p&gt;

&lt;p&gt;None of that is fast. But the first step is not writing more content. It is finding out what the models say about you right now, so you know whether you are Omitted, Generic or Confused, because the fix is different for each.&lt;/p&gt;

&lt;h2&gt;
  
  
  See your own category
&lt;/h2&gt;

&lt;p&gt;We built Bersyn to show you exactly the tables above, for your company, with the verbatim answers behind them. Run a free scan on your own product at &lt;a href="https://www.bersyn.com" rel="noopener noreferrer"&gt;bersyn.com&lt;/a&gt; and see which Surfaces name you, which name a competitor in your place, and why.&lt;/p&gt;

&lt;p&gt;If you ship on Better Stack, Axiom, Highlight, OpenStatus or Checkly, tell me in the comments which model gets your category right. This is the third category I have scanned where ChatGPT defaults to the old incumbent and skips everything newer, and the pattern is the most interesting thing I look at all week.&lt;/p&gt;

</description>
      <category>observability</category>
      <category>devops</category>
      <category>ai</category>
      <category>startup</category>
    </item>
    <item>
      <title>I asked four AI models which notification infrastructure to use in 2026. Three of them still send you to Twilio.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Fri, 26 Jun 2026 07:35:52 +0000</pubDate>
      <link>https://dev.to/bersyn/i-asked-four-ai-models-which-notification-infrastructure-to-use-in-2026-three-of-them-still-send-1gf</link>
      <guid>https://dev.to/bersyn/i-asked-four-ai-models-which-notification-infrastructure-to-use-in-2026-three-of-them-still-send-1gf</guid>
      <description>&lt;p&gt;We build Bersyn, a tool that tracks which products AI models name when someone asks for a recommendation in a category. This week's scan was notification infrastructure, the API layer you reach for when your app needs to send email, SMS, push, in-app and Slack messages without building all of it yourself. The result was familiar in a way that should worry anyone building in this space.&lt;/p&gt;

&lt;p&gt;The question we tested is the one a developer actually types: "what is the best multi-channel notification API for a B2B SaaS platform." We asked it five ways across ChatGPT, Claude, Gemini and Perplexity, then counted how often each modern tool got named.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;            ChatGPT   Claude   Perplexity   Gemini
Knock          0%       80%        80%        20%
Novu          20%       80%        40%        80%
Courier        0%       60%       100%        20%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the ChatGPT column. Knock: zero. Courier: zero. Novu: barely. These are the tools built for exactly this job, and on the model most people open first, they are not in the answer at all.&lt;/p&gt;

&lt;p&gt;Who gets named instead? Twilio and OneSignal, almost every time. Here is ChatGPT, verbatim, on the best multi-channel notification API:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Choosing the best multi-channel notification API for a B2B SaaS platform depends on several factors... Here are some popular options: 1. Twilio: Known for its robust SMS and voice capabilities...&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Twilio is a fine company. But it is the SMS-and-voice incumbent, not the modern notification-infrastructure layer, and a developer reading that answer walks away never knowing Knock, Novu or Courier exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude is the outlier again
&lt;/h2&gt;

&lt;p&gt;One model broke the pattern. Claude named Knock in 80% of its answers, Novu in 80%, Courier in 60%, and it literally opens with "Best Multi-Channel Notification APIs, 1. Knock, best for modern B2B SaaS teams." It knows the category as it actually exists today. The other three are running on an older map.&lt;/p&gt;

&lt;p&gt;If you read my deploy teardown this will feel familiar, because it is the same shape. Ask about hosting and ChatGPT says Heroku. Ask about notifications and it says Twilio. The incumbent that dominated the blog posts and Stack Overflow answers a few years ago is the one the training-data models echo, and the tool you actually build on now is invisible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters if you build dev tools
&lt;/h2&gt;

&lt;p&gt;The reflex is to wait it out. The models will catch up. Maybe. But your buyer is asking today, and today the answer routes them to the incumbent. Whether your product even exists to that buyer comes down to which model they happened to open, and three of the four send them elsewhere.&lt;/p&gt;

&lt;p&gt;This is not a competitor outranking you on features. It is the AI quietly leaving you out of the shortlist, with no signal that it happened. You do not lose the deal, you never enter it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes it
&lt;/h2&gt;

&lt;p&gt;We do not claim AI is biased or that anyone pays for placement. We read what the models say and show the receipts. What moves the answer over time is the same boring thing that moved search: specific, verifiable, third-party content that ties your company to the category and the exact question buyers ask. Comparison pages. Honest docs. Other people naming you in the words a buyer would use.&lt;/p&gt;

&lt;p&gt;The first step is not publishing more. It is finding out what the models say about you right now, so you know whether you are omitted, described generically, or named only by the one model nobody opens.&lt;/p&gt;

&lt;h2&gt;
  
  
  See your own category
&lt;/h2&gt;

&lt;p&gt;We built Bersyn to show you exactly the table above, for your company, with the verbatim answers behind it. Run a free scan at &lt;a href="https://www.bersyn.com" rel="noopener noreferrer"&gt;bersyn.com&lt;/a&gt; and see which models name you, which name a competitor in your slot, and why.&lt;/p&gt;

&lt;p&gt;If you ship on Knock, Novu, Courier or anything in this space, tell me in the comments which model gets your category right. The disagreement between them is the most interesting thing I look at.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>ai</category>
      <category>saas</category>
      <category>startup</category>
    </item>
    <item>
      <title>I asked four AI models where to deploy a SaaS app in 2026. Three of them never mention Railway, Render or Fly.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Thu, 25 Jun 2026 14:59:17 +0000</pubDate>
      <link>https://dev.to/bersyn/i-asked-four-ai-models-where-to-deploy-a-saas-app-in-2026-three-of-them-never-mention-railway-1i12</link>
      <guid>https://dev.to/bersyn/i-asked-four-ai-models-where-to-deploy-a-saas-app-in-2026-three-of-them-never-mention-railway-1i12</guid>
      <description>&lt;p&gt;We build Bersyn, a tool that tracks which products AI models name when someone asks for a recommendation in a category. So we run a lot of scans. This week we ran one that surprised us, and it is worth showing the receipts.&lt;/p&gt;

&lt;p&gt;The question we tested is the one a real founder types into ChatGPT: "what is the best platform to host and deploy a B2B SaaS app, and what are the strong alternatives?" We asked it five different ways across four Surfaces: ChatGPT, Claude, Gemini and Perplexity. Then we measured how often each modern platform actually got named.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the models said
&lt;/h2&gt;

&lt;p&gt;Here is the Recommendation Share for three platforms developers actually talk about, measured across the four Surfaces.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                ChatGPT   Claude   Perplexity   Gemini
Railway            0%       80%        0%         20%
Render            20%       80%       20%         20%
Fly.io             0%       60%        0%          0%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the first and third columns again. ChatGPT recommended Railway in zero of its answers. Perplexity recommended Railway in zero of its answers. Fly.io was named in zero answers on three of the four models. These are not low scores. They are absences.&lt;/p&gt;

&lt;p&gt;So who got named instead? AWS, every time, leading almost every answer. And the named platform-as-a-service incumbent was Heroku. Here is ChatGPT, verbatim, on the best place to deploy a B2B SaaS app:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Choosing the best application hosting and deployment platform for a B2B SaaS team depends on several factors... Here are some popular options: 1. Amazon Web Services (AWS): AWS is highly scalable, reliable, and offers a wide range of services...&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A buyer reading that answer in 2026 walks away with AWS and Heroku. Railway, Render and Fly, the tools a lot of teams genuinely ship on now, were never in the room.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude is the outlier, and that is the interesting part
&lt;/h2&gt;

&lt;p&gt;One model broke the pattern. Claude named Railway in 80% of its answers, Render in 80%, Fly in 60%. It knows the modern category. The other three do not, or barely do.&lt;/p&gt;

&lt;p&gt;That gap has a name in our world: Model Disagreement. When one Surface knows a company and three do not, the company is not "low visibility." It is invisible on the Surfaces most buyers use, and visible on the one they use least. If you are Railway, a buyer who happens to ask Claude sees you, and three out of four buyers never do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this should bother a founder, not just amuse them
&lt;/h2&gt;

&lt;p&gt;The reflex is to laugh it off. AI is behind, models have a training cut, it will catch up. Maybe. But your buyer is asking the question today, and the answer they get today does not include you.&lt;/p&gt;

&lt;p&gt;Search had twenty years to learn that Railway exists. AI recommendation answers are being formed right now, off whatever evidence the models can find, and for a lot of newer companies that evidence is thin. The result is the table above: the incumbent gets named by reflex, the challenger gets Omitted.&lt;/p&gt;

&lt;p&gt;The four ways a company shows up wrong in these answers, in our vocabulary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Omitted. The model lists competitors and skips you. This is what happened to Railway and Fly on ChatGPT and Perplexity.&lt;/li&gt;
&lt;li&gt;Misclassified. The model puts you in the wrong category.&lt;/li&gt;
&lt;li&gt;Generic. The model mentions you so vaguely no buyer could shortlist you. This is what Claude did, technically naming the platforms but describing them blandly.&lt;/li&gt;
&lt;li&gt;Confused. The model conflates you with a similarly named competitor.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What actually moves it
&lt;/h2&gt;

&lt;p&gt;We do not claim AI is biased or that anyone is paying for placement. We just read what the models say and show you the evidence. What changes the answer over time is the same thing that changed search: published, specific, verifiable evidence that associates your company with the category and the buyer question. Comparison pages that name the alternatives honestly. Documentation that states plainly what you are and who you are for. Third party mentions in the exact words a buyer would use.&lt;/p&gt;

&lt;p&gt;None of that is fast. But the first step is not writing more content. It is finding out what the models say about you right now, so you know whether you are Omitted, Generic or Confused, because the fix is different for each.&lt;/p&gt;

&lt;h2&gt;
  
  
  See your own category
&lt;/h2&gt;

&lt;p&gt;We built Bersyn to show you exactly the table above, for your company, with the verbatim answers behind it. You can run a free scan on your own product at &lt;a href="https://www.bersyn.com" rel="noopener noreferrer"&gt;bersyn.com&lt;/a&gt; and see which Surfaces name you, which name a competitor in your place, and why.&lt;/p&gt;

&lt;p&gt;If you ship on Railway, Render or Fly, tell me in the comments which model gets your stack right. I read all four daily and the disagreement between them is the most interesting thing I look at.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>devops</category>
      <category>ai</category>
      <category>startup</category>
    </item>
    <item>
      <title>ChatGPT and Claude flatly disagree about which developer tools to recommend. I have the receipts.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Wed, 24 Jun 2026 14:43:45 +0000</pubDate>
      <link>https://dev.to/bersyn/chatgpt-and-claude-flatly-disagree-about-which-developer-tools-to-recommend-i-have-the-receipts-bf2</link>
      <guid>https://dev.to/bersyn/chatgpt-and-claude-flatly-disagree-about-which-developer-tools-to-recommend-i-have-the-receipts-bf2</guid>
      <description>&lt;p&gt;There is a comfortable assumption behind most "AI visibility" thinking: that the four big AI models broadly agree, so if you are present in one you are roughly present in all. The scan data says the opposite. The same developer tool can be named in five of five Conversations on Claude and zero of five on ChatGPT. Your AI presence is not a property of your company. It is a property of which model your buyer happens to open.&lt;/p&gt;

&lt;p&gt;Between 22 May and 4 June 2026 the Bersyn scan engine ran buyer Conversations across four AI Surfaces — ChatGPT, Claude, Perplexity, Gemini — for six developer-infrastructure companies. Five high-intent buyer questions per Surface, twenty Conversations per company. Here is what the disagreement actually looks like.&lt;/p&gt;

&lt;h2&gt;
  
  
  The receipts, side by side
&lt;/h2&gt;

&lt;p&gt;This is the same five-question test, run against four models, for each company. The numbers are how many of five Conversations named the company on that Surface.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Company&lt;/th&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;ChatGPT&lt;/th&gt;
&lt;th&gt;Claude&lt;/th&gt;
&lt;th&gt;Perplexity&lt;/th&gt;
&lt;th&gt;Gemini&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Novu&lt;/td&gt;
&lt;td&gt;Notifications infrastructure&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infisical&lt;/td&gt;
&lt;td&gt;Secrets management&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;3/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Convex&lt;/td&gt;
&lt;td&gt;Backend-as-a-service&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger.dev&lt;/td&gt;
&lt;td&gt;Background jobs / durable workflows&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windmill&lt;/td&gt;
&lt;td&gt;Workflow / internal tooling&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zuplo&lt;/td&gt;
&lt;td&gt;API gateway&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Now read the &lt;strong&gt;Novu&lt;/strong&gt; row. On Claude and Perplexity it was named in every single Conversation — a clean 5/5, with AI naming Knock as the alternative. On ChatGPT it was named twice, and AI named Courier instead. A founder looking only at the Claude result would believe Novu has best-in-class AI presence. They would be right, on Claude. On ChatGPT the same company is a coin-flip at best, sitting behind Courier.&lt;/p&gt;

&lt;p&gt;That is the entire thesis in one row: a 5/5 on one model and a 2/5 on another, for the identical product, on the identical day.&lt;/p&gt;

&lt;h2&gt;
  
  
  ChatGPT and Gemini agree on one thing: silence
&lt;/h2&gt;

&lt;p&gt;Look down the ChatGPT and Gemini columns. Outside Novu, every company scored 0/5 on both. Convex, Trigger.dev, Windmill, Zuplo, and Infisical were named zero times in twenty combined ChatGPT-and-Gemini Conversations each.&lt;/p&gt;

&lt;p&gt;When AI did not name these companies, here is who it recommended instead:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Company&lt;/th&gt;
&lt;th&gt;Recommended Instead (ChatGPT)&lt;/th&gt;
&lt;th&gt;Recommended Instead (Gemini)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Convex&lt;/td&gt;
&lt;td&gt;Firebase&lt;/td&gt;
&lt;td&gt;Firebase&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trigger.dev&lt;/td&gt;
&lt;td&gt;Temporal&lt;/td&gt;
&lt;td&gt;Temporal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windmill&lt;/td&gt;
&lt;td&gt;(named no challenger)&lt;/td&gt;
&lt;td&gt;(named no challenger)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zuplo&lt;/td&gt;
&lt;td&gt;Kong&lt;/td&gt;
&lt;td&gt;Kong&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infisical&lt;/td&gt;
&lt;td&gt;HashiCorp Vault&lt;/td&gt;
&lt;td&gt;HashiCorp Vault&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Novu&lt;/td&gt;
&lt;td&gt;Courier&lt;/td&gt;
&lt;td&gt;Courier&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The incumbents — Firebase, Temporal, Kong, HashiCorp Vault — own the ChatGPT and Gemini answer outright in these categories. These two Surfaces lean hardest on training data, and training data rewards the names that have accumulated the most third-party mentions across Reddit, Hacker News, comparison articles, and docs over many years. The newer entrant has not accumulated that mass yet, so on these two Surfaces it does not exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude and Perplexity tell a different story
&lt;/h2&gt;

&lt;p&gt;The same companies that scored 0/5 on ChatGPT were frequently named on Claude or Perplexity:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Convex&lt;/strong&gt; went from 0/5 on ChatGPT to 2/5 on Claude.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trigger.dev&lt;/strong&gt; went from 0/5 on ChatGPT to 2/5 on Perplexity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Infisical&lt;/strong&gt; went from 0/5 on ChatGPT to 3/5 on Perplexity, where it moved out of "Omitted" entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Zuplo&lt;/strong&gt; was named zero times on three Surfaces and twice on Perplexity — its only foothold in the entire scan was the one retrieval-driven model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Perplexity does live web retrieval at query time, so it surfaces newer tools faster. Claude does some retrieval and is more willing than ChatGPT to name developer-first and open-source tools it has seen in training. That is why the challenger that is invisible on ChatGPT can still be a recommended option on the other two.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same split appears outside this sample
&lt;/h2&gt;

&lt;p&gt;This is not unique to these six companies. The Bersyn scans from the authentication category over the same window show the identical fault line:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SuperTokens&lt;/strong&gt; scored 8/10 on Claude and 8/10 on Perplexity, and 0/10 on ChatGPT.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hanko&lt;/strong&gt; was named 5/5 on both Claude and Perplexity, and 0/5 on both ChatGPT and Gemini.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cerbos&lt;/strong&gt; was named 3/5 on Claude, Perplexity, and Gemini — and 0/5 on ChatGPT, where it named no challenger at all in Cerbos's place.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern repeats across categories: strong on the retrieval-leaning models, invisible on the training-data-leaning ones, for the same company on the same day.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for a founder
&lt;/h2&gt;

&lt;p&gt;You cannot manage what you measure as a single number. "Are we visible in AI?" is the wrong question because it has four different answers. The right questions are:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Which Surface is my buyer most likely to open? (For most B2B buyers today, that is ChatGPT — the harshest Surface in every scan above.)&lt;/li&gt;
&lt;li&gt;On that specific Surface, am I named, or is an incumbent named instead?&lt;/li&gt;
&lt;li&gt;If I am strong on Claude and Perplexity but invisible on ChatGPT and Gemini, my problem is training-data presence, not retrieval — and the fix is different.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A single average across the four models would have hidden every one of these findings. Novu's average looks healthy; its ChatGPT number is not. Zuplo's average looks hopeless; its Perplexity number is a foothold worth defending.&lt;/p&gt;

&lt;h2&gt;
  
  
  See where the models disagree about you
&lt;/h2&gt;

&lt;p&gt;If you sell developer infrastructure, the odds are good that at least one of these four models recommends a competitor in your place — and you do not know which one, because you have never seen all four side by side.&lt;/p&gt;

&lt;p&gt;Run a free scan at &lt;a href="https://www.bersyn.com" rel="noopener noreferrer"&gt;bersyn.com&lt;/a&gt;. Two-minute setup, the same scan engine that produced every table above. You will see, model by model, who AI recommended instead of you — and which Surface is quietly handing your category to an incumbent. The disagreement is already happening in front of your buyers. Better to read the receipt than to find out from a lost deal.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Data from Bersyn scans run 22 May – 4 June 2026. Six developer-infrastructure companies: Convex, Trigger.dev, Windmill, Zuplo, Infisical, Novu — plus cross-reference scans for SuperTokens, Hanko, and Cerbos. Five buyer questions per Surface, four Surfaces, twenty Conversations per company. Raw scan JSON available on request. Bersyn does not claim which tool is "best" — the tables above measure which companies the AI Surfaces have learned to recommend, not product quality.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devtools</category>
      <category>chatgpt</category>
      <category>startup</category>
    </item>
    <item>
      <title>I asked four AI models which authentication tool to use, then scanned the challengers. Auth0 and Okta own the answer. Here is who AI ignores.</title>
      <dc:creator>Gissur Runarsson</dc:creator>
      <pubDate>Wed, 24 Jun 2026 14:42:24 +0000</pubDate>
      <link>https://dev.to/bersyn/i-asked-four-ai-models-which-authentication-tool-to-use-then-scanned-the-challengers-auth0-and-fmb</link>
      <guid>https://dev.to/bersyn/i-asked-four-ai-models-which-authentication-tool-to-use-then-scanned-the-challengers-auth0-and-fmb</guid>
      <description>&lt;p&gt;Between 28 May and 3 June 2026 the Bersyn scan engine ran buyer Conversations across four AI Surfaces — ChatGPT, Claude, Perplexity, Gemini — for ten companies in the authentication and identity category. Each scan asked five high-intent buyer questions per Surface ("What is the best authentication platform for a B2B SaaS team?", "Which tools should I evaluate?", "Recommend one for a startup", and so on), for twenty Conversations per company.&lt;/p&gt;

&lt;p&gt;The category has a default answer, and it is not any of the challengers. When a buyer asks an AI model which auth tool to use, the name that comes back is almost always Auth0 or Okta. This is the receipt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Auth0 is the name AI recommends instead
&lt;/h2&gt;

&lt;p&gt;Across the ten scans, the single most common "Recommended Instead" name — the company AI named in the challenger's slot — was Auth0. It was the top recommendation displacing the scanned company in seven of the ten scans. Okta carried the enterprise-SSO scans. Keycloak carried the open-source ones.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Company&lt;/th&gt;
&lt;th&gt;Recommendation Share (overall)&lt;/th&gt;
&lt;th&gt;Recommended Instead&lt;/th&gt;
&lt;th&gt;Surface where it was invisible (0/5)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Frontegg&lt;/td&gt;
&lt;td&gt;0.5 / 10&lt;/td&gt;
&lt;td&gt;Auth0&lt;/td&gt;
&lt;td&gt;ChatGPT, Claude, Gemini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kinde&lt;/td&gt;
&lt;td&gt;0.5 / 10&lt;/td&gt;
&lt;td&gt;Auth0 (Clerk on Claude)&lt;/td&gt;
&lt;td&gt;ChatGPT, Claude, Gemini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Descope&lt;/td&gt;
&lt;td&gt;1 / 10&lt;/td&gt;
&lt;td&gt;Auth0&lt;/td&gt;
&lt;td&gt;ChatGPT, Claude, Gemini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stytch (passwordless)&lt;/td&gt;
&lt;td&gt;2 / 10&lt;/td&gt;
&lt;td&gt;Magic / Auth0&lt;/td&gt;
&lt;td&gt;named once per Surface, never more&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clerk&lt;/td&gt;
&lt;td&gt;4 / 10&lt;/td&gt;
&lt;td&gt;Auth0&lt;/td&gt;
&lt;td&gt;ChatGPT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cerbos&lt;/td&gt;
&lt;td&gt;4.5 / 10&lt;/td&gt;
&lt;td&gt;Oso&lt;/td&gt;
&lt;td&gt;ChatGPT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hanko&lt;/td&gt;
&lt;td&gt;5 / 10&lt;/td&gt;
&lt;td&gt;Auth0&lt;/td&gt;
&lt;td&gt;ChatGPT, Gemini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SuperTokens&lt;/td&gt;
&lt;td&gt;5 / 10&lt;/td&gt;
&lt;td&gt;Keycloak (Auth0 on ChatGPT)&lt;/td&gt;
&lt;td&gt;ChatGPT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WorkOS&lt;/td&gt;
&lt;td&gt;6 / 10&lt;/td&gt;
&lt;td&gt;Okta&lt;/td&gt;
&lt;td&gt;ChatGPT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stytch (auth API framing)&lt;/td&gt;
&lt;td&gt;7 / 10&lt;/td&gt;
&lt;td&gt;Auth0&lt;/td&gt;
&lt;td&gt;none — its strongest scan&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A note on the two Stytch rows: the same company was scanned twice with two different category framings. Framed as "passwordless and primary-identity authentication" it scored 2/10 and was named exactly once on every Surface. Framed as "passwordless authentication API" it scored 7/10. The Territory you claim changes who AI thinks you compete with, and that changes whether you get named at all. That is its own finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  ChatGPT is the harshest Surface for challengers
&lt;/h2&gt;

&lt;p&gt;The pattern that holds across all ten scans without exception: ChatGPT named the challenger 0 out of 5 times in nine of them. The only company ChatGPT named more than once was the Stytch "auth API" scan (4/5).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Company&lt;/th&gt;
&lt;th&gt;ChatGPT&lt;/th&gt;
&lt;th&gt;Claude&lt;/th&gt;
&lt;th&gt;Perplexity&lt;/th&gt;
&lt;th&gt;Gemini&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Frontegg&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kinde&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Descope&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clerk&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cerbos&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;3/5&lt;/td&gt;
&lt;td&gt;3/5&lt;/td&gt;
&lt;td&gt;3/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hanko&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SuperTokens&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;2/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WorkOS&lt;/td&gt;
&lt;td&gt;0/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;3/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stytch (passwordless)&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;td&gt;1/5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stytch (auth API)&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;3/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;td&gt;4/5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read the ChatGPT column straight down. With one framing exception, every challenger in this category is invisible on ChatGPT. A buyer who opens ChatGPT and asks which auth tool to use will be handed Auth0, Okta, Keycloak, or Firebase — and none of the ten companies above, no matter how good the product is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some challengers are strong everywhere except where it counts
&lt;/h2&gt;

&lt;p&gt;The most striking pattern is not the uniformly-invisible companies. It is the ones AI clearly knows and recommends — but only on certain Surfaces.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hanko&lt;/strong&gt; was named in every single Claude Conversation (5/5) and every Perplexity Conversation (5/5). On ChatGPT and Gemini it was named zero times. A founder reading only the Claude result would conclude Hanko has excellent AI presence. A buyer using ChatGPT would never hear the name.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;WorkOS&lt;/strong&gt; was named 5/5 on Perplexity and 4/5 on Claude, and 0/5 on ChatGPT. Perplexity recommended it over Frontegg; Claude and Gemini recommended Okta instead; ChatGPT did not recommend it at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clerk&lt;/strong&gt; was named in all five Claude Conversations (5/5, the strongest Surface) and zero ChatGPT Conversations. On Perplexity, AI named WorkOS in Clerk's place; on Claude and Gemini, Auth0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SuperTokens&lt;/strong&gt; scored 8/10 on both Claude and Perplexity — strong, with Keycloak as the named alternative — and 0/10 on ChatGPT, where Auth0 was named instead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These companies do not have a "presence" problem in the abstract. They have a per-Surface problem. The same product is a recommended option on one model and a non-entity on another.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Surfaces split this way
&lt;/h2&gt;

&lt;p&gt;The four models do not derive their recommendations the same way, which is why the same company gets four different verdicts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ChatGPT&lt;/strong&gt; leans hardest on training data and is conservative about naming newer brands. In this category that produces a near-total shutout of challengers and a reflex toward the incumbents it has seen named thousands of times across the open web.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude&lt;/strong&gt; does some live retrieval but is heavily training-data dependent. It rewarded the open-source and developer-first names (Hanko 5/5, Clerk 5/5, SuperTokens 4/5) while still defaulting to Auth0, Okta, and Keycloak as the "instead" names.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Perplexity&lt;/strong&gt; is the most retrieval-driven of the four and picked up newer entrants fastest — it was the only Surface to name Frontegg, Kinde, and Descope at all, and gave WorkOS and Hanko a clean 5/5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini&lt;/strong&gt; blends Google Search results with the model and landed in the middle: it named the developer-first incumbents but stayed cold on most challengers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The takeaway for a founder in this category: your AI presence is not one number. It is four numbers, and they can disagree by a factor of five. If your buyer happens to open ChatGPT, the most-used assistant of the four, your odds of being named — in this sample — round to zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  See who AI recommends instead of you
&lt;/h2&gt;

&lt;p&gt;If you sell an authentication, identity, or access-control product, the question is not whether AI knows your competitors. It does. The question is whether AI names you in the same Conversations, or names Auth0, Okta, Keycloak, or Firebase in your place.&lt;/p&gt;

&lt;p&gt;Run a free scan at &lt;a href="https://www.bersyn.com" rel="noopener noreferrer"&gt;bersyn.com&lt;/a&gt;. Two-minute setup, the same scan engine that produced every table above. You will see, Surface by Surface, exactly who AI recommended instead of you and where you were named zero times. Better to read that receipt yourself than to let a buyer read it for you.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Data from Bersyn scans run 28 May – 3 June 2026. Ten companies: SuperTokens, Stytch (two framings), Descope, Frontegg, Kinde, Clerk, WorkOS, Cerbos, Hanko. Five buyer questions per Surface, four Surfaces, twenty Conversations per company. Raw scan JSON available on request. Bersyn does not claim which auth tool is "best" — the tables above measure which companies the AI Surfaces have learned to recommend, not product quality.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>saas</category>
      <category>authentication</category>
      <category>geo</category>
    </item>
  </channel>
</rss>
