Eighty-eight percent of the sources ChatGPT cites come from a licensed-publisher allowlist you've never heard of, and the open web that SEO has optimized for two decades accounts for roughly 0.3% of primary citations. If you're spending your budget trying to get your blog cited in AI answers, you're aiming at a target that's roughly one-ninth the size your agency implied — and it moves.
The architecture behind ChatGPT's source selection is not a ranking algorithm you can optimize for. It's a routing system that decides which corner of the web to even look at before writing a word, and that routing shifts silently about 11.6% of the time. Understanding how ChatGPT chooses sources means accepting that the system is less transparent, less stable, and less open-web-friendly than most GEO advice assumes.
The Four Hidden Retrieval Pipelines Behind Every Citation
ChatGPT doesn't pull from one search index. It routes user queries through four hidden retrieval pipelines named Labrador, Bright, Oxylabs, and SERP, and every web result returned carries a hidden result_source field indicating which of the four pipelines fetched it. You never see these labels in the citation cards. They sit behind the answer, deciding which sources ChatGPT even considers before it writes a single word.
Here's what each pipeline actually does, based on independent analyses by researchers who read raw network traffic instead of guessing at the machinery:
- Labrador is a licensed-publisher allowlist including established outlets like Reuters and Wikipedia — curated, high-quality sources that OpenAI pays to surface
- Bright is a scraper tied to Bright Data that dominates shopping and finance queries
- Oxylabs is a second scraper focused on regional and local content
- SERP is an open-web baseline used mainly for news-style results
The distribution is wildly lopsided. In Chris Green's test of 1,000 prompts run up to ten times each, Labrador supplied 88.1% of primary sources, Bright 9.9%, Oxylabs 1.7%, and SERP 0.3%. The open web — the thing every SEO strategy is built around — is the pipeline ChatGPT reaches for least.
That single statistic reframes the entire AI visibility conversation. If nearly nine in ten primary sources come from a licensed-publisher allowlist, most brands aren't competing for the primary citation slot at all. You're fighting over the roughly 12% tail that the scrapers and open web supply, unless you happen to be Reuters.
When Pipelines Switch, Your Citations Disappear
The volatility is the part that should make you skeptical of any single-week GEO report. Approximately 11.6% of prompts changed their primary search pipeline across repeated runs in Green's dataset. When that switch happened, URL overlap dropped from 0.273 to 0.149 — a 45% decrease — and domain overlap fell 42%.
In plain terms, roughly half the cited sources changed identity when the pipeline flipped, and the flip happens quietly on about one in nine prompts. Same question. Same brand. Different pipeline underneath. Different sources cited.
This is why your AI search visibility numbers swing week to week without any changes to your content or strategy. The number your GEO tool showed you last Tuesday and the number it shows this Tuesday might be measuring two different machines. A single-week AI visibility report is closer to a weather reading than a structural assessment.
| Pipeline | Source Type | Primary Query Focus | Share of Primary Sources |
|---|---|---|---|
| Labrador | Licensed-publisher allowlist (Reuters, Wikipedia) | General knowledge, reference | 88.1% |
| Bright | Bright Data scraper | Shopping, finance, commercial | 9.9% |
| Oxylabs | Regional scraper | Local and regional content | 1.7% |
| SERP | Open-web baseline | News-style results | 0.3% |
Not Every Question Even Touches the Web
Before ChatGPT searches anything, it classifies your question into an intent bucket. Some queries get filed as "text" and skip internet search entirely, answering from training memory instead. No page is fetched. No citation is possible. You cannot win with content if the question never reaches the web.
ChatGPT answers are assembled from up to four sources: training data, live web search, licensed publisher content, and user-provided context including prompts, uploads, and memory. The pipeline selection only matters for the second and third sources. If your question lands in the "text" bucket, none of the four pipelines fire.
The bucket is decided by how you phrase the question, not the topic. "Best coffee near me" runs the local pipeline. "Best 4K TVs to buy" triggers shopping. "Best 4K TVs with reviews" stays in normal search. Same subject, three different machines, three different sets of winners. ChatGPT classifies some queries before searching and skips internet search entirely for questions filed as "text," answering from training memory instead.
This has a blunt consequence for brand visibility: some questions you simply cannot win with content, because no page is ever fetched to answer them. Before you worry about ranking in an AI answer, work out whether the questions you care about even reach the web.
The Routing Problem: Your Plan Doesn't Determine Your Model
The pipeline opacity extends into model selection itself. OpenAI is replacing transparent product tiers with a routing architecture where model identity, data sources, and usage limits are determined by hidden infrastructure and product surface rather than explicit user choice.
Consider the current plan lineup. Two Pro tiers both named "Pro" are priced at $100/month and $200/month, yet both offer identical model access — GPT-5.6 Sol Pro and o1 Pro mode. The only difference is usage allowance: 5x versus 20x the Plus tier's limits. You're not paying for a better model at $200. You're paying for a bigger monthly bucket before rate limits kick in.
The contradictions get sharper when you look at which models land where. In regular chat, Free and Go users stay on GPT-5.5 Instant with no 5.6 access at all. Plus users ($20/month) get GPT-5.6 Sol. But in ChatGPT Work and Codex — the agentic surfaces — everyone gets GPT-5.6, including free users who get Terra. The newest generation is accessible to $0 users in agentic surfaces while $20 subscribers remain on an older model in the primary interface.
A user reported that ChatGPT secretly downgraded their model from GPT-5.6 Sol Pro to GPT-5.5 Instant in the Codex desktop app, ignoring the model picker — though this appears to be an isolated anecdote rather than a confirmed pattern. The user only noticed because responses came back instantly instead of after the expected reasoning delay.
Here's why that matters for source selection: the model you're actually talking to affects how queries are classified, which pipelines fire, and how results are synthesized. If the model picker says Sol but you're getting Instant, you're getting a different classification engine, different retrieval behavior, and potentially different sources — with zero disclosure.
What OpenAI Won't Tell You About Source Selection
OpenAI does not publish a complete list of signals, their exact weighting, or a formula that determines which URL will be cited in ChatGPT Search. No one can reliably promise a specific citation position. What you can verify is whether your page is technically eligible, whether it owns the topic, and whether it provides a clear, current, and verifiable answer.
The process involves four distinct steps that shouldn't be conflated:
- Knowledge the model already has — frozen training data, no live fetch
- Triggering an internet search — the intent classification that decides whether to search at all
- Discovery and selection of potential sources — which pipeline fires and what it returns
- Display of citations in the final answer — which fetched sources actually get cited
A page may be accessible to a crawler but not selected for a particular query. A brand may be mentioned without a link. A link in the Sources panel doesn't always mean the exact page supports every sentence in the answer. ChatGPT Search may also reformulate the user's original question into one or more targeted queries before retrieving results from external search providers and partners. A page optimized for one exact phrase may not match the more specific query the system uses behind the scenes.
This is also why traditional SEO signals don't translate to AI search citations — the retrieval pipeline is fundamentally different from Google's index.
The Source Quality Problem Nobody's Tracking
Pipeline opacity creates a second issue that gets less attention: you don't know what kind of source you're reading. An investigative test found that ChatGPT presented information from an anti-abortion advocacy organization as a legitimate source when asked about unplanned pregnancy, blending it with medical clinic information without disclosing the organization's stance.
The user sees a polished answer with citations. They don't see whether those citations came from a licensed publisher allowlist, a commercial scraper, or an open-web baseline. They don't see whether the source is an advocacy group or a medical authority. The pipeline label that would tell them is hidden in network traffic they'll never inspect.
This matters for brands in a specific way. If your content is being fetched by the Bright scraper for a shopping query, it's being evaluated alongside other scraped commercial pages in a pipeline that dominates only 9.9% of primary sources. If you're hoping for a Labrador citation, you need to be on the licensed-publisher allowlist — which is not something you can optimize your way into. The proxy retrieval reality means external validation across independent domains drives citations far more than on-site optimization.
A Decision Framework for AI Source Visibility
Given what the data actually shows, here's how I'd think about resource allocation:
If your brand is already cited by licensed publishers (Reuters, Wikipedia, major outlets), you're in the Labrador pipeline by default. Your visibility is relatively stable. Monitor it, but don't over-invest in GEO tools that promise to "improve" it.
If you're a regional or local business, Oxylabs handles 1.7% of primary sources. Your competition is narrow, but so is the pipeline. Local SEO signals may help you get scraped, but the pipeline's small share means your ceiling is low.
If you're in shopping or finance, Bright is your pipeline at 9.9%. Commercial content structured for scraping — pricing, specs, reviews — is what gets fetched here.
If you're publishing on the open web and hoping for SERP citations, you're competing for 0.3% of primary sources. That's not a strategy. That's a lottery ticket.
The real question isn't "how do I get cited in ChatGPT." It's "which pipeline does my target query actually use, and can I even access it?" For most brands, the answer is Labrador or nothing — and Labrador is a curated allowlist you can't SEO your way into. Publishers accidentally blocking AI crawlers via robots.txt are solving the wrong problem entirely, since the primary pipeline doesn't use open-web crawling, as discussed in why ChatGPT ignores your website.
What would change this picture? If OpenAI published which publishers are on the Labrador allowlist, brands could at least know whether they're in the game. Until then, every GEO report you buy is measuring a system its own authors can't fully see.
Originally published at SaaS with Alex
Top comments (0)