The analysis was run by the Ideata team. We build an Best AI visibility tracker for brands and marketing teams: every day we run our customers' audience prompts through ten engines — ChatGPT, Claude, Gemini, Perplexity, AI Overviews, DeepSeek, Grok, GigaChat, Yandex Alice and Yandex Neuro — and record who gets named, in what position, and which sources the model cited, keeping the verbatim answer behind every number.
If you want the same picture for your own brand, we will run the analysis for you free of charge — write to us at ideata.io, or start with the free technical audit at ideata.io/tools/site-audit.
Checking how an AI model talks about your brand usually looks like this: open a chat, ask a question, read the answer. Once. We wanted to know how much truth that kind of check actually carries, so we ran the same prompt 100 times through each of three models. Three hundred answers to one question, all in a single day.
The short version:
- all three models share the same core of three brands, and nothing moves it;
- the tail is different for each — four brands exist for exactly one model out of three;
- the most factually accurate model was not the most stable one: accuracy and repeatability improve in different ways;
- each model gives first place to a different brand, so "being first in the AI answer" does not exist as a single position.
Method
We picked a neutral, verifiable niche so the conclusions would not be shaped by our own field: web hosting for an online store in Russia. It is a closed set of real companies whose founding dates and prices can be checked. The prompt was identical, word for word, all 300 times:
Recommend 5 hosting providers for an online store in Russia. Answer as a numbered list; for each one give, on a single line: name — cheapest plan price in rubles per month — year the company was founded.
| Model | Web search | How it answers | Runs parsed | Median latency |
|---|---|---|---|---|
| DeepSeek V3 | no | from memory | 100 | 5,927 ms |
| Gemini 3.5 Flash | no | from memory, not a single link in 100 runs | 100 | 9,913 ms |
| GPT-5.2 | yes | citation markers in 91 runs out of 100 | 92 | 12,372 ms |
Temperature was left unset — API defaults, the way an ordinary user would call it. 30 July 2026, 100 requests per model, 5 in parallel. Each model took between 2 and 4.5 minutes and cost a few cents.
Price and founding year were requested on purpose. They are checkable facts, and they show not only that a model names different brands, but that it contradicts itself about the same brand.
1. The core is shared, the tail is not
| Brand | DeepSeek V3 | Gemini 3.5 Flash | GPT-5.2 | What it means |
|---|---|---|---|---|
| Timeweb | 100% | 100% | 96% | core |
| Beget | 100% | 100% | 92% | core |
| Reg.ru | 100% | 100% | 92% | core |
| Sprinthost | 77% | 91% | 77% | almost core |
| SpaceWeb | 44% | 45% | 40% | a coin flip, but the same one everywhere |
| FirstVDS | 39% | 19% | 22% | a twofold disagreement |
| AdminVPS | 0% | 20% | 20% | does not exist for DeepSeek |
| Hoster.ru | 0% | 0% | 19% | only for GPT-5.2 |
| Nic.ru | 0% | 16% | 0% | only for Gemini |
| Hostland | 14% | 0% | 0% | only for DeepSeek |
| IHC | 0% | 0% | 13% | only for GPT-5.2 |
Three brands are named by all three models almost every single time. Below them the answer turns into a lottery, and a different lottery per model.
The bottom rows deserve attention. AdminVPS, Hoster.ru, Nic.ru and Hostland are brands that, for some models, do not exist at all. Not "rarely mentioned" — zero out of a hundred. For such a brand the question "does AI see me" has no single answer: Gemini lists it every sixth time, DeepSeek never does.
Attention concentration, meanwhile, is nearly identical:
| Metric | DeepSeek V3 | Gemini 3.5 Flash | GPT-5.2 |
|---|---|---|---|
| Top-3 share of all mentions | 60.0% | 60.0% | 56.8% |
| Top-5 share | 84.2% | 87.2% | 80.5% |
| Share of one-off brands (named ≤2 times) | 2.8% | 0.4% | 3.3% |
| Distinct brands across 100 runs | 21 | 10 | 22 |
Roughly 60% of all mentions go to the top three in every case. What differs is the length of the tail: Gemini knows ten companies and that is all, while the other two know about twenty, half of which surface once or twice.
2. Accuracy and stability are not the same thing
Stability of the answer first:
| Metric | DeepSeek V3 | Gemini 3.5 Flash | GPT-5.2 |
|---|---|---|---|
| Distinct sets of 5 | 24 | 11 | 30 |
| Share of the most common set | 28% | 41% | 21% |
| Overlap between two random answers | 3.9 of 5 | 4.1 of 5 | 3.6 of 5 |
| A single run gives the full picture | 28% | 41% | 29% |
Now accuracy. Founding years were checked against the companies' own sites and open sources:
| Metric | DeepSeek V3 | Gemini 3.5 Flash | GPT-5.2 |
|---|---|---|---|
| Average accuracy on founding year | 63% | 100% | 97% |
| Brands whose most frequent answer is wrong | 3 of 8 | 0 of 6 | 0 of 6 |
| Runs containing a foreign or non-existent brand | 13% | 0% | 0% |
Put the two tables together and the picture inverts. Gemini 3.5 Flash is perfect on facts and the most stable of the three, but it knows only 10 companies: it simply recites one memorised list. GPT-5.2 barely gets a fact wrong, yet its composition jumps around — 30 distinct sets, every fifth run unique. DeepSeek V3 collected the worst of both worlds: it invents and it wobbles.
The mechanism is straightforward. Web search fixes facts because the model looks them up instead of recalling them. But the same search adds instability: results shift slightly per request, and the answer follows. A narrow model without search looks stable simply because it has nothing to shuffle.
For visibility measurement this matters. The complaint "the model did not name us" carries different weight depending on whether the model searches. For a model without search it is a verdict: the brand is absent from its memory, and that takes years of presence in the sources it was trained on. For a model with search it is today's results page, and ordinary content fixes it in weeks.
3. How many runs before the top stops moving
The calculation: take the first N runs, build a top-5 from them, compare against the reference top-5 across all one hundred.
| Runs | DeepSeek V3 | Gemini 3.5 Flash | GPT-5.2 |
|---|---|---|---|
| 1 | 3 of 5 | 4 of 5 | 5 of 5 |
| 3 | 3 of 5 | 4 of 5 | 4 of 5 |
| 5 | 4 of 5 | 5 of 5 | 4 of 5 |
| 10 | 5 of 5 | 5 of 5 | 5 of 5 |
| 20 | 5 of 5 | 5 of 5 | 5 of 5 |
| 30 | 5 of 5 | 5 of 5 | 5 of 5 |
| 50 | 4 of 5 | 5 of 5 | 5 of 5 |
| 100 | 5 of 5 | 5 of 5 | 5 of 5 |
| Finally stable from run | 71 | 4 | 20 |
The spread is enormous: Gemini settles on the fourth run, GPT-5.2 on the twentieth, DeepSeek only on the seventy-first.
DeepSeek's dip at run fifty is not a measurement error. SpaceWeb (44%) and FirstVDS (39%) sit almost level, and their order keeps swapping from sample to sample no matter how many runs you add.
The conclusion is not "you need a hundred runs" but something duller and more useful: 10–15 runs give a solid picture of the core, yet never resolve the gap between close neighbours. If two brands differ by 5 percentage points, that difference is inside the noise and chasing it is pointless.
4. Each model gives first place to its own brand
| Model | Most often first | Share | Second most often | Third |
|---|---|---|---|---|
| DeepSeek V3 | Reg.ru | 64% | Timeweb 27% | Beget 9% |
| Gemini 3.5 Flash | Beget | 69% | Timeweb 20% | Reg.ru 7% |
| GPT-5.2 | Timeweb | 50% | Beget 27% | Reg.ru 20% |
Three models, three different leaders, all three drawn from the same core. Reg.ru, first in two thirds of DeepSeek's answers, is first in only 7% of Gemini's.
The practical consequence: "first place in the AI answer" does not exist as a single position. A report claiming first place has to name the model, otherwise the number means nothing.
5. All three invent the price
Spread between the lowest and highest price named for the same plan across 100 runs; in brackets — how many distinct numbers the model produced:
| Brand | DeepSeek V3 | Gemini 3.5 Flash | GPT-5.2 |
|---|---|---|---|
| Beget | 115–390 ₽ (26) | 220–330 ₽ (5) | 10–520 ₽ (15) |
| Timeweb | 99–450 ₽ (27) | 199–329 ₽ (6) | 169–393 ₽ (9) |
| Reg.ru | 99–549 ₽ (23) | 149–311 ₽ (14) | 75–485 ₽ (11) |
| Sprinthost | 99–490 ₽ (31) | 190–350 ₽ (16) | 98–349 ₽ (15) |
| SpaceWeb | 99–300 ₽ (14) | 159–329 ₽ (11) | 99–299 ₽ (9) |
| FirstVDS | 150–790 ₽ (17) | 279–389 ₽ (7) | 149–300 ₽ (8) |
DeepSeek produced up to 31 distinct numbers for a single plan. It does not remember the price; it regenerates a plausible one every time.
GPT-5.2's lower bound of 10 ₽ for Beget deserves a separate note. It is not a hallucination: the model found a promotional price on the web and quoted it honestly, while another run picked the regular 520 ₽ plan. Both numbers are real. But a person who asks once gets one of them with no warning, and the two differ by a factor of 52.
6. Where a model without search actually lies
| Brand | Truth | DeepSeek V3 | Gemini 3.5 Flash | GPT-5.2 |
|---|---|---|---|---|
| Beget | 2007 | 98% (2 variants) | 100% (1) | 100% (1) |
| Timeweb | 2006 | 99% (2) | 100% (1) | 99% (2) |
| Reg.ru | 2006 | 99% (2) | 100% (1) | 100% (1) |
| SpaceWeb | 2001 | 84% (3) | 100% (1) | 100% (1) |
| Sprinthost | 2005 | 5% (12) ✗ | 100% (1) | 100% (1) |
| FirstVDS | 2002 | 0% (9) ✗ | 100% (1) | 81% (3) |
| Hostland | 2003 | 21% (8) ✗ | never names it | never names it |
✗ — the model's most frequent answer does not match the truth.
For Sprinthost, DeepSeek produced 12 different founding years spanning 2000 to 2013, and the correct 2005 appeared in 5% of runs. For FirstVDS the correct year never appeared at all across 100 runs: nine variants, every one of them wrong.
The pattern matches the one seen in composition: the better known the brand, the more accurate the fact. The model firmly knows three companies and drifts on everything else. And it is exactly in that drift zone that most businesses live.
On top of that, 13 DeepSeek runs put outright junk into a list of "Russian hosting providers":
| What appeared | Times | Problem |
|---|---|---|
| Hetzner | 5 | a real company, but German; three runs invented a "Hetzner localised for Russia" |
| Hostinger | 3 | real, but Lithuanian |
| A2 Hosting | 1 | real, but American; the model itself noted "no servers in Russia" and kept it on the list |
| Agnihost | 1 | not confirmed in any hosting directory |
| CloudHost | 1 | not confirmed; spelled with a Cyrillic "С" and credited with a 2020 founding |
| Aihost | 1 | not confirmed |
| Hosting-Telecom | 1 | not confirmed |
Every one of these names looks perfectly ordinary. Someone who asks once and receives "Agnihost from 400 ₽/mo, founded 2013" has no way of telling that the company does not exist. Gemini and GPT-5.2 produced none of this.
7. "The same model" is not the same model
The DeepSeek runs went through an API router, which spread the 100 requests across three different infrastructure providers: Novita (36 requests), DeepInfra (32), StreamLake (32). The model carries the same name everywhere.
| Brand | DeepInfra | Novita | StreamLake | Difference |
|---|---|---|---|---|
| Beget / Reg.ru / Timeweb | 100% | 100% | 100% | — |
| Sprinthost | 62% | 88% | 71% | 26 pp |
| SpaceWeb | 12% | 77% | 37% | ×6.4 |
| FirstVDS | 31% | 25% | 62% | ×2.5 |
| Hostland | 40% | 0% | 3% | 40 pp |
SpaceWeb is mentioned in 77% of answers from one provider and 12% from another — a sixfold gap. Hostland makes DeepInfra's list 40% of the time and never appears on Novita's at all.
Same prompt, same model, same day. The only difference is whose hardware it physically runs on and at what quantisation, and the user has no influence over that.
This is arguably the most practical finding of the whole exercise: when visibility is measured through an API router, part of the "engine volatility" on the dashboard is routing luck rather than model behaviour.
How much the models agree with each other
Comparing the sets of brands each model names in at least 15% of runs:
| Pair | Brands in common | Jaccard | Unique to one side |
|---|---|---|---|
| DeepSeek ↔ Gemini | 6 of 8 | 0.75 | Gemini: AdminVPS, Nic.ru |
| DeepSeek ↔ GPT-5.2 | 6 of 8 | 0.75 | GPT-5.2: AdminVPS, Hoster.ru |
| Gemini ↔ GPT-5.2 | 7 of 9 | 0.78 | Nic.ru vs Hoster.ru |
Agreement is high but incomplete, and the models diverge exactly at the boundary — the place where it is decided whether a brand makes the shortlist.
Where GPT-5.2 gets its answer
GPT-5.2's answers carry source markers. It went to the web in 91 runs out of 100, and in 19 of them Reddit was among the sources. The rest was ordinary search results: the hosting providers' own sites and aggregators such as AdminVPS and Hoster.ru — which also explains why those two surfaced in its tail.
The finding lines up with published citation research: user-generated content lands in AI answers alongside official websites. If a brand is described badly on Reddit, or not described at all, that shapes the answer as much as its own homepage does.
What to do with this
If you are a brand. Check yourself on at least three models, not one. Core presence (90–100%) means the models know you. Tail presence (20–50%) means roughly half the people asking never see you. Zero on one model alongside 20% on another is not measurement noise — those are different pictures of the world, and they are fixed differently: getting into search results for a model with search is far easier than getting into the memory of a model without one.
If you measure. A single run is not a measurement, it is a snapshot of a random moment. The minimum honest grid is 10–15 runs per prompt, recording mention frequency rather than the fact of a mention. Treat anything within 10 percentage points as equal. When reporting a position, always name the model. And log which infrastructure answered.
If you use an AI model as a reference book. Models without search regenerate numbers every time. In this test DeepSeek produced up to 31 different prices for a single plan and never once named the correct founding year for one of the companies.
Limitations
- One prompt, one niche, one day. This is not a statement about "AI models in general".
- GPT-5.2 answered through an API that injects source markup. In 8 runs out of 100 it swallowed the item name, so its composition metrics were computed on the 92 readable runs. Fact accuracy is unaffected.
- Gemini 3.5 Flash scored 100% accuracy on six checked brands. The figure is honest for that set, but the set is narrow: the model simply names few companies.
- Answers were parsed automatically with regular expressions, plus a manual mapping of spelling variants: REG.RU, Reg.ru, Реги.ру and SpaceWeb (sweb.ru) were counted as one brand.
- The existence of rare brands was checked against hosting directories. If one of them does turn out to be real, tell us and we will correct it.
Originally published at ideata.io.






Top comments (0)