DEV Community

Searchless
Searchless

Posted on • Originally published at searchless.ai

ChatGPT's Citation Inequality: Why Travel Gets 22% and Education Gets 5%

Originally published on The Searchless Journal

ChatGPT does not cite sources equally. According to Similarweb's 2026 Generative AI Landscape report, ChatGPT includes citations in 22.6% of travel-related answers but only 4.8% of education-related answers. That is a 4.7x gap between the most-cited and least-cited industries. The overall citation rate across all topics is 6.8%, meaning most verticals sit well below the level needed for meaningful referral traffic from AI-generated answers.

This is not a minor variation. It is structural discovery inequality. Brands in travel, retail, and sports operate in citation-rich environments where investment in Generative Engine Optimization produces direct, measurable returns. Brands in education, health, and media operate in citation-poor environments where the same investment yields a fraction of the visibility. The implication is uncomfortable but urgent: GEO ROI is not uniform. It is vertical-dependent. And most brands are investing blind, without knowing whether their industry's citation baseline gives them a fighting chance.

The data comes from Similarweb's analysis of millions of ChatGPT conversations, tracking when and how the AI includes source links in its responses. Search Engine Journal reported the findings on July 27, marking the first time topic-level citation behavior has been quantified at scale. The numbers change how brands should think about AI visibility strategy. A travel company investing in GEO can expect roughly one in four answers to include a citation. An education company investing the same amount can expect fewer than one in twenty. The playing field is not level, and pretending otherwise wastes budget.

The Numbers: Citation Rate by Industry

The Similarweb data is stark in its spread. Here is how ChatGPT's citation rate breaks down by topic.

Travel leads all categories at 22.6%. When users ask ChatGPT about destinations, flights, hotels, itineraries, or travel advice, the AI includes source links in nearly one-quarter of its answers. Retail follows at 13.5%. Sports sits at 10.7%. Finance answers include citations 8.0% of the time. The overall average across all topics is 6.8%.

Below the average, the drop-off is steep. Technology answers cite sources less often than the average. Health, media, and entertainment all fall below the mean. Education sits at the bottom with 4.8%. For context, a 4.8% citation rate means that out of every 100 education-related questions asked on ChatGPT, approximately 95 receive answers with no link to any external source whatsoever.

The practical consequence is immediate. A brand in the travel space optimizing for ChatGPT visibility has a structural advantage that a brand in education cannot match through effort alone. No amount of schema markup, answer-first content structuring, or LLMs.txt implementation will close a 4.7x gap that is built into how the model handles different topics.

This is not a critique of ChatGPT's design. The variation likely reflects underlying training data density, the nature of user queries in each vertical, and the model's assessment of when source attribution adds value. Travel questions naturally involve specific, time-sensitive information (prices, schedules, reviews) where citing a source improves answer quality. Education questions often involve conceptual explanations where the model can synthesize from its training data without needing to point users elsewhere. The mechanism is understandable. The strategic consequence is what matters.

The Growth Story: 1.3% to 6.8% in Eleven Months

The headline citation rate of 6.8% represents significant growth. In June 2025, ChatGPT's overall citation rate was approximately 1.3%. Over eleven months, that rate increased roughly fivefold. This is one of the clearest signals that AI search is evolving toward a citation-inclusive model rather than a closed-loop answer engine.

But the growth has not been linear. Similarweb's longitudinal data shows that the citation rate was highly volatile throughout late 2025 and early 2026. The rate first crossed 6% in October 2025, then dropped to approximately 4.5% by February 2026, before climbing again through the spring to reach the current 6.8% mark.

This volatility matters for brands tracking their AI visibility. A brand that measured its ChatGPT citation rate in January 2026 and concluded GEO was not working may have been measuring during a trough. A brand that measured in October 2025 and saw strong results may have been capturing a peak that was not sustainable. The lesson is that single-point-in-time AI visibility measurement is unreliable. Citation rates move with model updates, product changes, and behavioral shifts. Continuous monitoring is the only way to distinguish signal from noise.

The Resoneo data reinforces this point. In April 2026, Search Engine Journal reported that Resoneo observed ChatGPT citing approximately 20% fewer websites per response after the GPT-5.3 Instant update. A single model update compressed the citation surface by one-fifth. Brands that were being cited before the update may have disappeared from answers overnight, with no notification, no explanation, and no recourse. This is the citation volatility problem, and it makes static GEO audits insufficient.

Source Type Preferences: Not Just Whether You Get Cited, but Who Gets Cited

The Similarweb data also reveals that ChatGPT's choice of source types varies dramatically by topic. This adds a second dimension to citation inequality. It is not just that some industries get cited more often. It is that the kinds of sources ChatGPT prefers change depending on what the user is asking about.

For beauty and personal care queries, 54.7% of cited sources are retail and e-commerce sites. This means a beauty brand that sells direct-to-consumer through its own online store has a structural citation advantage over a beauty brand that relies on third-party retailers, editorial coverage, or social media presence. ChatGPT is looking for commerce-ready product pages to cite, not brand awareness pieces.

For travel queries, 54.1% of cited sources are reviews and user-generated content. TripAdvisor, Yelp, Google Reviews, and similar platforms dominate ChatGPT's travel citations. A hotel with thousands of positive reviews on TripAdvisor is more likely to be cited than a hotel with a beautifully optimized website but thin review presence. The citation economy rewards platforms where users contribute structured evaluations, not brands that publish the most polished content.

For finance queries, the picture shifts again. Finance answers cite finance-specific sites 36.6% of the time and news publishers 28.0% of the time. This means a fintech company hoping to be cited needs presence on finance-specific platforms (investopedia, NerdWallet, Bankrate) and in reputable financial news outlets. A fintech company that publishes only on its own blog, no matter how well-structured, is playing on a field where ChatGPT prefers established financial authorities.

Across all topics, the overall source type breakdown is: reviews and user-generated content at 28.9%, news publishers at 26.0%, and retail and e-commerce sites at 14.1%. The remaining 31% is distributed across blogs, educational sites, corporate pages, and other categories.

The strategic implication is that source-type alignment matters as much as citation rate. A brand in a vertical where ChatGPT prefers user-generated content needs a review strategy, not just a content strategy. A brand in a vertical where ChatGPT prefers news publishers needs a PR and earned media strategy. A brand in a vertical where ChatGPT prefers e-commerce sites needs a product page optimization strategy. Generic GEO advice ("write answer-first content, add schema markup") is insufficient when the model's source preferences vary so dramatically by topic.

The Vertical Strategy Split: What Citation-Poor Industries Should Do

For brands in citation-rich verticals like travel, retail, and sports, the path is relatively straightforward. Invest in the content structures, platform presence, and technical accessibility that maximize citation probability. Ensure product pages, review profiles, and category-leading content are crawlable by GPTBot. Monitor citation rates continuously. Optimize for the source types ChatGPT prefers in your vertical.

For brands in citation-poor verticals like education, health, and media, the strategy must be fundamentally different. Chasing citations that statistically will not appear is a waste of resources. The overall citation rate of 6.8% means that most ChatGPT answers are self-contained, and in citation-poor verticals, the self-containment rate approaches 95%.

ChatGPT citation benchmark by source type and answer engine

Three alternative strategies exist for citation-poor verticals.

First, entity optimization. Even when ChatGPT does not cite a source, it draws on training data and retrieved context to generate answers. If your brand is a recognized entity in the model's knowledge base with consistent attributes (founding date, location, key personnel, product categories, notable achievements), you can influence answers without being cited. Entity optimization means ensuring your brand's structured data, Wikipedia presence, Wikidata entries, and consistent web mentions give the model accurate, authoritative information to draw from. This is invisible visibility. The user sees your brand mentioned in the answer, but there is no link to click. For education brands, health information providers, and media companies, entity optimization may deliver more practical visibility than chasing citations.

Second, direct LLM influence through training data. Brands in citation-poor verticals can focus on being present in the data that trains and grounds language models. This means publishing authoritative, factually dense, well-structured content that gets crawled and indexed by AI training pipelines. The goal is not to earn a citation in a specific answer but to shape the model's understanding of your brand, your category, and your expertise. This is a longer-term play with less measurable immediate impact, but it addresses the reality that most AI answers are synthesized from training data, not from real-time citations.

Third, structured data and knowledge graph presence. Schema markup, Wikidata entries, and knowledge graph connections give AI models machine-readable context about your brand. While structured data alone will not generate a citation where the model has decided not to cite, it ensures that when your brand does appear in an answer, the information is accurate, complete, and properly attributed. For health and education brands where misinformation carries high stakes, structured data presence is a defensive necessity, not just an offensive GEO tactic.

Why This Data Changes the GEO Conversation

The GEO industry has operated on an implicit assumption: if you optimize your content correctly, you can earn visibility in AI search results regardless of your industry. The Similarweb data breaks this assumption. Citation rate is not primarily a function of content quality or optimization effort. It is a function of topic-level model behavior that varies by a factor of nearly five.

This does not mean GEO is irrelevant for citation-poor verticals. It means GEO strategies must be calibrated to vertical realities. A travel brand and an education brand cannot run the same GEO playbook and expect comparable results. The travel brand's GEO investment will produce measurable citation growth, referral traffic, and attributable conversions. The education brand's GEO investment will produce invisible influence, entity recognition, and brand presence in synthesized answers without links. Both have value. But they are different values, and they require different measurement frameworks.

The GEO conversation needs to mature beyond "are you cited?" to "how does your industry's citation baseline shape your AI visibility strategy?" A travel brand with a 22.6% citation baseline should measure citation rate growth, referral traffic from ChatGPT, and conversion attribution from AI-referred visitors. An education brand with a 4.8% citation baseline should measure brand mention accuracy, entity recognition consistency, and share of voice in synthesized answers. These require different tools, different methodologies, and different expectations.

This is also why AI visibility audits matter more than ever. Running a comprehensive audit across multiple engines, hundreds of queries, and multiple verticals gives brands the baseline data needed to set realistic expectations and choose the right strategy. Without knowing your industry's citation baseline and your position within it, GEO investment is a gamble. With that data, it becomes a calculated investment with a defensible expected return.

The Volatility Factor: Model Updates Can Reshape Citations Overnight

The Similarweb data provides a snapshot. But the citation landscape is not static. It shifts with every model update, product change, and behavioral adjustment ChatGPT makes.

The Resoneo finding from April 2026 is the clearest example. After the GPT-5.3 Instant update, ChatGPT cited approximately 20% fewer websites per response. A single model update wiped out one-fifth of the citation surface. Brands that had built their AI visibility strategy on citation presence suddenly found themselves invisible. The volatility was not announced, not explained, and not reversed.

This volatility has a compounding effect on citation inequality. When ChatGPT reduces its citation surface, it does not do so uniformly. Citation-rich verticals like travel may absorb the reduction and still maintain high rates. Citation-poor verticals may drop below measurable thresholds. A 20% reduction on a 22.6% base leaves travel at approximately 18%. A 20% reduction on a 4.8% base leaves education at approximately 3.8%. The absolute gap narrows slightly, but the relative impact on citation-poor verticals is more severe because they had less margin to lose.

Brands need to build volatility into their AI visibility planning. Single-point measurements are unreliable. Month-over-month tracking is the minimum viable monitoring frequency. Brands should expect citation rate swings of 20% or more following major model updates and should have contingency strategies for citation droughts, just as they would for search algorithm fluctuations.

Beyond ChatGPT: The Multi-Engine Question

The Similarweb data covers ChatGPT specifically. But ChatGPT is one of several AI engines where brands need visibility. Perplexity, Google AI Overviews, Gemini, and Claude all have different citation behaviors.

Perplexity's evidence-first design produces citation rates that are structurally higher than ChatGPT's. Because Perplexity was built around source transparency from the beginning, it treats citations as a core product feature rather than an optional enhancement. Growth Memo data suggests Perplexity generates 2-3x higher click-through rate per citation compared to Google AI Overviews, because users on Perplexity are conditioned to click through to sources.

Google AI Overviews operates within the traditional SERP, where citations compete with ads, organic results, and other SERP features for user attention. Gemini's citation behavior is evolving as Google integrates AI Mode into its billion-user product portfolio.

The implication is that ChatGPT's citation inequality by industry likely exists across other engines too, though the specific rates and rankings may differ. A travel brand that is well-cited on ChatGPT is also likely to perform well on Perplexity and Google AI Overviews, because the underlying factors (time-sensitive information, review density, product-specific pages) that make travel citation-rich on ChatGPT probably apply elsewhere. An education brand that struggles on ChatGPT will likely struggle on other engines too, though perhaps not as severely on Perplexity given its evidence-first design.

Multi-engine monitoring is essential. Brands should not assume that citation performance on one engine predicts performance on all engines. But they should expect that the structural patterns (travel and retail citing more, education and health citing less) are broadly consistent across engines, because they reflect underlying characteristics of how AI models handle different types of information.

If your brand needs to understand its citation baseline across ChatGPT, Perplexity, Gemini, and Claude, run a comprehensive AI visibility audit. It is the only way to know whether your industry gives you a citation-rich or citation-poor starting point.

Sources

  • Similarweb, 2026 Generative AI Landscape Report (citation rates by topic, source type preferences, longitudinal citation data)
  • Search Engine Journal, Matt Southern, "ChatGPT Links Out Most on Travel Queries, Data Shows" (July 27, 2026)
  • Search Engine Journal, Roger Monttti, AI citation pattern analysis across engines (May 2026)
  • Resoneo / Search Engine Journal, ChatGPT citation contraction after GPT-5.3 Instant update (April 2026)
  • Growth Memo, AI Mode user behavior study (CTR differentials by engine)
  • Alphabet Q2 2026 Earnings Release (Google AI Mode billion-user portfolio, Gemini 950M users)

FAQ

What is a good ChatGPT citation rate for my industry?

There is no universal "good" rate. The Similarweb data shows that citation rates range from 22.6% (travel) to 4.8% (education). A travel brand cited in 18% of relevant queries is underperforming its industry baseline. An education brand cited in 6% of relevant queries is outperforming its industry baseline by 25%. Benchmark against your vertical, not against the cross-industry average.

If my industry has a low citation rate, should I still invest in GEO?

Yes, but differently. In citation-poor verticals, focus on entity optimization (ensuring your brand is accurately represented in AI training data and knowledge graphs), structured data presence, and brand mention accuracy rather than chasing citation links that are statistically unlikely to appear. The value of GEO in citation-poor verticals is invisible influence rather than measurable referral traffic.

How often do ChatGPT citation rates change?

Citation rates fluctuate continuously, with significant shifts following model updates. The GPT-5.3 Instant update in April 2026 reduced citation surface by approximately 20%. The overall rate has swung between 4.5% and 6.8% over the past eleven months. Monthly monitoring is the minimum viable tracking frequency.

Does citation inequality exist on other AI engines too?

Likely yes, though the specific rates differ. Perplexity's evidence-first design produces structurally higher citation rates across all topics. Google AI Overviews operates within the SERP context where citations compete with other elements. The underlying pattern (information-seeking topics cite more, conceptual topics cite less) is likely consistent across engines because it reflects how language models handle different query types.


Want to know what your brand's AI visibility looks like across ChatGPT, Perplexity, Gemini, and Claude? Run a free AI visibility audit and get your industry-specific citation baseline.

Or explore Searchless pricing and service options for comprehensive AI visibility management.

Top comments (0)