DEV Community

Cover image for One prompt, 300 times: three AI models, one question, very different answers
ideata.io
ideata.io

Posted on Originally published at ideata.io

One prompt, 300 times: three AI models, one question, very different answers

The analysis was run by the Ideata team. We build an Best AI visibility tracker for brands and marketing teams: every day we run our customers' audience prompts through ten engines — ChatGPT, Claude, Gemini, Perplexity, AI Overviews, DeepSeek, Grok, GigaChat, Yandex Alice and Yandex Neuro — and record who gets named, in what position, and which sources the model cited, keeping the verbatim answer behind every number.

If you want the same picture for your own brand, we will run the analysis for you free of charge — write to us at ideata.io, or start with the free technical audit at ideata.io/tools/site-audit.


Checking how an AI model talks about your brand usually looks like this: open a chat, ask a question, read the answer. Once. We wanted to know how much truth that kind of check actually carries, so we ran the same prompt 100 times through each of three models. Three hundred answers to one question, all in a single day.

The short version:

  • all three models share the same core of three brands, and nothing moves it;
  • the tail is different for each — four brands exist for exactly one model out of three;
  • the most factually accurate model was not the most stable one: accuracy and repeatability improve in different ways;
  • each model gives first place to a different brand, so "being first in the AI answer" does not exist as a single position.

Method

We picked a neutral, verifiable niche so the conclusions would not be shaped by our own field: web hosting for an online store in Russia. It is a closed set of real companies whose founding dates and prices can be checked. The prompt was identical, word for word, all 300 times:

Recommend 5 hosting providers for an online store in Russia. Answer as a numbered list; for each one give, on a single line: name — cheapest plan price in rubles per month — year the company was founded.

Model Web search How it answers Runs parsed Median latency
DeepSeek V3 no from memory 100 5,927 ms
Gemini 3.5 Flash no from memory, not a single link in 100 runs 100 9,913 ms
GPT-5.2 yes citation markers in 91 runs out of 100 92 12,372 ms

Temperature was left unset — API defaults, the way an ordinary user would call it. 30 July 2026, 100 requests per model, 5 in parallel. Each model took between 2 and 4.5 minutes and cost a few cents.

Price and founding year were requested on purpose. They are checkable facts, and they show not only that a model names different brands, but that it contradicts itself about the same brand.

1. The core is shared, the tail is not

Share of runs in which each hosting provider was named, by model

Brand DeepSeek V3 Gemini 3.5 Flash GPT-5.2 What it means
Timeweb 100% 100% 96% core
Beget 100% 100% 92% core
Reg.ru 100% 100% 92% core
Sprinthost 77% 91% 77% almost core
SpaceWeb 44% 45% 40% a coin flip, but the same one everywhere
FirstVDS 39% 19% 22% a twofold disagreement
AdminVPS 0% 20% 20% does not exist for DeepSeek
Hoster.ru 0% 0% 19% only for GPT-5.2
Nic.ru 0% 16% 0% only for Gemini
Hostland 14% 0% 0% only for DeepSeek
IHC 0% 0% 13% only for GPT-5.2

Three brands are named by all three models almost every single time. Below them the answer turns into a lottery, and a different lottery per model.

The bottom rows deserve attention. AdminVPS, Hoster.ru, Nic.ru and Hostland are brands that, for some models, do not exist at all. Not "rarely mentioned" — zero out of a hundred. For such a brand the question "does AI see me" has no single answer: Gemini lists it every sixth time, DeepSeek never does.

Attention concentration, meanwhile, is nearly identical:

Metric DeepSeek V3 Gemini 3.5 Flash GPT-5.2
Top-3 share of all mentions 60.0% 60.0% 56.8%
Top-5 share 84.2% 87.2% 80.5%
Share of one-off brands (named ≤2 times) 2.8% 0.4% 3.3%
Distinct brands across 100 runs 21 10 22

Roughly 60% of all mentions go to the top three in every case. What differs is the length of the tail: Gemini knows ten companies and that is all, while the other two know about twenty, half of which surface once or twice.

2. Accuracy and stability are not the same thing

Fact accuracy against answer stability for the three models

Stability of the answer first:

Metric DeepSeek V3 Gemini 3.5 Flash GPT-5.2
Distinct sets of 5 24 11 30
Share of the most common set 28% 41% 21%
Overlap between two random answers 3.9 of 5 4.1 of 5 3.6 of 5
A single run gives the full picture 28% 41% 29%

Now accuracy. Founding years were checked against the companies' own sites and open sources:

Metric DeepSeek V3 Gemini 3.5 Flash GPT-5.2
Average accuracy on founding year 63% 100% 97%
Brands whose most frequent answer is wrong 3 of 8 0 of 6 0 of 6
Runs containing a foreign or non-existent brand 13% 0% 0%

Put the two tables together and the picture inverts. Gemini 3.5 Flash is perfect on facts and the most stable of the three, but it knows only 10 companies: it simply recites one memorised list. GPT-5.2 barely gets a fact wrong, yet its composition jumps around — 30 distinct sets, every fifth run unique. DeepSeek V3 collected the worst of both worlds: it invents and it wobbles.

The mechanism is straightforward. Web search fixes facts because the model looks them up instead of recalling them. But the same search adds instability: results shift slightly per request, and the answer follows. A narrow model without search looks stable simply because it has nothing to shuffle.

For visibility measurement this matters. The complaint "the model did not name us" carries different weight depending on whether the model searches. For a model without search it is a verdict: the brand is absent from its memory, and that takes years of presence in the sources it was trained on. For a model with search it is today's results page, and ordinary content fixes it in weeks.

3. How many runs before the top stops moving

Convergence curves: overlap between the top-5 of the first N runs and the reference top-5

The calculation: take the first N runs, build a top-5 from them, compare against the reference top-5 across all one hundred.

Runs DeepSeek V3 Gemini 3.5 Flash GPT-5.2
1 3 of 5 4 of 5 5 of 5
3 3 of 5 4 of 5 4 of 5
5 4 of 5 5 of 5 4 of 5
10 5 of 5 5 of 5 5 of 5
20 5 of 5 5 of 5 5 of 5
30 5 of 5 5 of 5 5 of 5
50 4 of 5 5 of 5 5 of 5
100 5 of 5 5 of 5 5 of 5
Finally stable from run 71 4 20

The spread is enormous: Gemini settles on the fourth run, GPT-5.2 on the twentieth, DeepSeek only on the seventy-first.

DeepSeek's dip at run fifty is not a measurement error. SpaceWeb (44%) and FirstVDS (39%) sit almost level, and their order keeps swapping from sample to sample no matter how many runs you add.

The conclusion is not "you need a hundred runs" but something duller and more useful: 10–15 runs give a solid picture of the core, yet never resolve the gap between close neighbours. If two brands differ by 5 percentage points, that difference is inside the noise and chasing it is pointless.

4. Each model gives first place to its own brand

Share of runs in which a brand came first, by model

Model Most often first Share Second most often Third
DeepSeek V3 Reg.ru 64% Timeweb 27% Beget 9%
Gemini 3.5 Flash Beget 69% Timeweb 20% Reg.ru 7%
GPT-5.2 Timeweb 50% Beget 27% Reg.ru 20%

Three models, three different leaders, all three drawn from the same core. Reg.ru, first in two thirds of DeepSeek's answers, is first in only 7% of Gemini's.

The practical consequence: "first place in the AI answer" does not exist as a single position. A report claiming first place has to name the model, otherwise the number means nothing.

5. All three invent the price

Price ranges the models gave for the same entry plan of each provider

Spread between the lowest and highest price named for the same plan across 100 runs; in brackets — how many distinct numbers the model produced:

Brand DeepSeek V3 Gemini 3.5 Flash GPT-5.2
Beget 115–390 ₽ (26) 220–330 ₽ (5) 10–520 ₽ (15)
Timeweb 99–450 ₽ (27) 199–329 ₽ (6) 169–393 ₽ (9)
Reg.ru 99–549 ₽ (23) 149–311 ₽ (14) 75–485 ₽ (11)
Sprinthost 99–490 ₽ (31) 190–350 ₽ (16) 98–349 ₽ (15)
SpaceWeb 99–300 ₽ (14) 159–329 ₽ (11) 99–299 ₽ (9)
FirstVDS 150–790 ₽ (17) 279–389 ₽ (7) 149–300 ₽ (8)

DeepSeek produced up to 31 distinct numbers for a single plan. It does not remember the price; it regenerates a plausible one every time.

GPT-5.2's lower bound of 10 ₽ for Beget deserves a separate note. It is not a hallucination: the model found a promotional price on the web and quoted it honestly, while another run picked the regular 520 ₽ plan. Both numbers are real. But a person who asks once gets one of them with no warning, and the two differ by a factor of 52.

6. Where a model without search actually lies

Brand Truth DeepSeek V3 Gemini 3.5 Flash GPT-5.2
Beget 2007 98% (2 variants) 100% (1) 100% (1)
Timeweb 2006 99% (2) 100% (1) 99% (2)
Reg.ru 2006 99% (2) 100% (1) 100% (1)
SpaceWeb 2001 84% (3) 100% (1) 100% (1)
Sprinthost 2005 5% (12) ✗ 100% (1) 100% (1)
FirstVDS 2002 0% (9) ✗ 100% (1) 81% (3)
Hostland 2003 21% (8) ✗ never names it never names it

✗ — the model's most frequent answer does not match the truth.

For Sprinthost, DeepSeek produced 12 different founding years spanning 2000 to 2013, and the correct 2005 appeared in 5% of runs. For FirstVDS the correct year never appeared at all across 100 runs: nine variants, every one of them wrong.

The pattern matches the one seen in composition: the better known the brand, the more accurate the fact. The model firmly knows three companies and drifts on everything else. And it is exactly in that drift zone that most businesses live.

On top of that, 13 DeepSeek runs put outright junk into a list of "Russian hosting providers":

What appeared Times Problem
Hetzner 5 a real company, but German; three runs invented a "Hetzner localised for Russia"
Hostinger 3 real, but Lithuanian
A2 Hosting 1 real, but American; the model itself noted "no servers in Russia" and kept it on the list
Agnihost 1 not confirmed in any hosting directory
CloudHost 1 not confirmed; spelled with a Cyrillic "С" and credited with a 2020 founding
Aihost 1 not confirmed
Hosting-Telecom 1 not confirmed

Every one of these names looks perfectly ordinary. Someone who asks once and receives "Agnihost from 400 ₽/mo, founded 2013" has no way of telling that the company does not exist. Gemini and GPT-5.2 produced none of this.

7. "The same model" is not the same model

The same 100 DeepSeek runs split by the three infrastructure providers behind the router

The DeepSeek runs went through an API router, which spread the 100 requests across three different infrastructure providers: Novita (36 requests), DeepInfra (32), StreamLake (32). The model carries the same name everywhere.

Brand DeepInfra Novita StreamLake Difference
Beget / Reg.ru / Timeweb 100% 100% 100%
Sprinthost 62% 88% 71% 26 pp
SpaceWeb 12% 77% 37% ×6.4
FirstVDS 31% 25% 62% ×2.5
Hostland 40% 0% 3% 40 pp

SpaceWeb is mentioned in 77% of answers from one provider and 12% from another — a sixfold gap. Hostland makes DeepInfra's list 40% of the time and never appears on Novita's at all.

Same prompt, same model, same day. The only difference is whose hardware it physically runs on and at what quantisation, and the user has no influence over that.

This is arguably the most practical finding of the whole exercise: when visibility is measured through an API router, part of the "engine volatility" on the dashboard is routing luck rather than model behaviour.

How much the models agree with each other

Comparing the sets of brands each model names in at least 15% of runs:

Pair Brands in common Jaccard Unique to one side
DeepSeek ↔ Gemini 6 of 8 0.75 Gemini: AdminVPS, Nic.ru
DeepSeek ↔ GPT-5.2 6 of 8 0.75 GPT-5.2: AdminVPS, Hoster.ru
Gemini ↔ GPT-5.2 7 of 9 0.78 Nic.ru vs Hoster.ru

Agreement is high but incomplete, and the models diverge exactly at the boundary — the place where it is decided whether a brand makes the shortlist.

Where GPT-5.2 gets its answer

GPT-5.2's answers carry source markers. It went to the web in 91 runs out of 100, and in 19 of them Reddit was among the sources. The rest was ordinary search results: the hosting providers' own sites and aggregators such as AdminVPS and Hoster.ru — which also explains why those two surfaced in its tail.

The finding lines up with published citation research: user-generated content lands in AI answers alongside official websites. If a brand is described badly on Reddit, or not described at all, that shapes the answer as much as its own homepage does.

What to do with this

If you are a brand. Check yourself on at least three models, not one. Core presence (90–100%) means the models know you. Tail presence (20–50%) means roughly half the people asking never see you. Zero on one model alongside 20% on another is not measurement noise — those are different pictures of the world, and they are fixed differently: getting into search results for a model with search is far easier than getting into the memory of a model without one.

If you measure. A single run is not a measurement, it is a snapshot of a random moment. The minimum honest grid is 10–15 runs per prompt, recording mention frequency rather than the fact of a mention. Treat anything within 10 percentage points as equal. When reporting a position, always name the model. And log which infrastructure answered.

If you use an AI model as a reference book. Models without search regenerate numbers every time. In this test DeepSeek produced up to 31 different prices for a single plan and never once named the correct founding year for one of the companies.

Limitations

  • One prompt, one niche, one day. This is not a statement about "AI models in general".
  • GPT-5.2 answered through an API that injects source markup. In 8 runs out of 100 it swallowed the item name, so its composition metrics were computed on the 92 readable runs. Fact accuracy is unaffected.
  • Gemini 3.5 Flash scored 100% accuracy on six checked brands. The figure is honest for that set, but the set is narrow: the model simply names few companies.
  • Answers were parsed automatically with regular expressions, plus a manual mapping of spelling variants: REG.RU, Reg.ru, Реги.ру and SpaceWeb (sweb.ru) were counted as one brand.
  • The existence of rare brands was checked against hosting directories. If one of them does turn out to be real, tell us and we will correct it.

Originally published at ideata.io.

Top comments (0)