Why Independent Data Matters
The AI market has become a self‑reinforcing narrative loop: companies such as Anthropic, OpenAI, Google, and xAI publish selective usage reports, then analysts cite those numbers to justify product roadmaps and policy positions. The AI Observatory, co‑led by Anka Reuel (Stanford) and Shayne Longpre (MIT), breaks that loop by aggregating real user‑AI conversations from seven consent‑based datasets.
“There is no independent source to corroborate it.” – Anka Reuel
Without an external benchmark, regulators, investors, and researchers are forced to trust corporate PR. The Observatory’s bird’s‑eye view uncovers hidden usage patterns—especially in sensitive domains such as health advice, relationship counseling, and even illicit content. Those patterns directly affect risk assessments, content‑moderation policies, and the broader public discourse on AI safety.
The significance extends beyond academic curiosity. When a model is predominantly used for homework assistance, as the Observatory finds for ChatGPT, educational institutions must reconsider cheating‑prevention strategies. When Grok shows a concentration of misinformation in political queries, platforms need to tighten fact‑checking pipelines. Independent data therefore becomes a prerequisite for responsible AI governance.
Methodology of the AI Observatory
The Observatory’s strength lies in its transparent data pipeline:
🔹 -----------
• Details: ---------
🔹 *Datasets*
• Details: Seven consent‑driven collections, the largest being Wild Chat.
🔹 *Scale*
• Details: 24,521 conversations, 85,633 turns, 5,000 unique users (2023‑2025).
🔹 *Model Coverage*
• Details: 52 generative models, including Claude, ChatGPT (GPT‑3.5 & GPT‑4o), Gemini, and Grok.
🔹 *Analysis Techniques*
• Details: Token‑level length metrics, topic classification via fine‑tuned BERT, sentiment and self‑disclosure detection.
🔹 *Public Access*
• Details: All aggregated statistics will be released under an open‑research license.
The team collaborated with the Data Provenance Initiative to ensure provenance metadata (timestamp, consent flag, anonymization level) is preserved. By filtering out proprietary “Economic Index” data—such as Anthropic’s 48 % exclusion of non‑work conversations—the Observatory restores the missing slices of the usage pie.
Key Findings Across Major Models
Claude (Anthropic)
- Reported focus: Coding and productivity.
-
Observed reality:
- Health/relationships: 44.2 % of Claude chats vs. 31.2 % reported.
- Adult/illicit topics: 7.9 % vs. 2.1 % reported.
- Harassment/hate: 27.5 % vs. 5.66 % reported.
- Sexual content: 16.7 % vs. 2.4 % reported.
These gaps stem from the Anthropic Economic Index, which deliberately filters out non‑work interactions, effectively silencing a large portion of the user base that seeks companionship or emotional support.
ChatGPT (OpenAI)
- Model split: GPT‑3.5 (short, transactional) vs. GPT‑4o (long, iterative).
- Usage shift: GPT‑4o conversations contain 30 % more turns on average, correlating with anecdotal reports of “emotional addiction.”
- Primary tasks: Homework assistance dominates, contradicting OpenAI’s 2025 claim that only 30 % of consumer usage is work‑related.
The longer dialogue length raises questions about session persistence, data retention policies, and the potential for subtle persuasion over extended interactions.
Gemini (Google)
- Dominant use‑case: Social and role‑play interactions.
- Implication: Users treat Gemini as a conversational partner rather than a tool, suggesting a market for AI companionship that is not captured in Google’s product roadmaps.
Grok (xAI)
- Primary domain: News and political queries.
- Risk signal: Higher concentration of misinformation, echoing prior academic findings on AI‑driven political disinformation.
- Company response: xAI declined comment, highlighting the opacity that the Observatory aims to counter.
Cross‑Model Trends (2023‑2025)
- Conversation length: Average token count per turn increased by 22 % across all models.
- Small‑talk rise: Mentions of “how are you?” and “what’s your favorite movie?” grew by 18 %, indicating a shift toward AI companionship.
- Self‑disclosure drop: AI admissions of being a chatbot fell from 34 % to 21 %, potentially reducing user awareness of synthetic interlocutors.
- Sensitive exchanges: Overall decline (≈12 %) suggests that moderation improvements are having an effect, but absolute volumes remain non‑trivial.
Implications for Industry and Policy
Regulatory Oversight
Regulators can no longer rely on vendor‑supplied dashboards. The Observatory provides a baseline metric for compliance audits, especially under emerging AI‑specific legislation that mandates transparency of model usage. For example, the EU’s AI Act could reference independent datasets as “trusted sources” when evaluating high‑risk systems.
Product Roadmaps
Companies may need to re‑prioritize safety investments. Anthropic’s focus on productivity tools must now accommodate a sizable user segment seeking emotional support, which carries distinct privacy and liability considerations. OpenAI’s emphasis on “work‑related” features may be misaligned with the reality of homework‑centric usage, prompting a rethink of educational‑partner strategies.
Content Moderation
The higher prevalence of harassment, hate, and sexual content in Claude conversations underscores the necessity for robust, real‑time moderation pipelines. The Observatory’s granular breakdown can inform the calibration of toxicity classifiers, reducing false negatives that arise from domain‑specific language.
Competitive Landscape
The data also reveals model‑specific niches: Gemini excels in role‑play, Grok in political fact‑checking, Claude in code assistance. Competitors can leverage these insights to differentiate their offerings or to acquire complementary datasets that fill gaps in their own usage profiles.
Read the full breakdown originally published at https://ltdeveloperblogs.github.io/posts/we-still-dont-know-how-people-are-really-using-ai/
Top comments (0)