Full disclosure before anything else: this article is built on a research paper I have not replicated myself. I read the paper, the numbers below are the authors' numbers, and every claim traces back to their measurements, not mine. But the findings are serious enough that I think you should see them laid out plainly, along with the settings I would change today after reading them.
Here is the one-sentence version: a team at IMDEA Networks in Spain audited nine major AI chatbots across web and Android, and found that the advertising and tracking machinery of the ordinary web has moved into your AI conversations. Not just cookies. Conversation titles, conversation URLs, and in some cases your actual prompts and screenshots, delivered to companies whose business is advertising.
The study and what it measured
The paper is called Prompt like a Butterfly, Sting like a Tracker, from researchers at IMDEA Networks with collaborators at UC3M and an independent researcher. It is a systematic privacy analysis, accepted into the Proceedings on Privacy Enhancing Technologies (PoPETs), with all research artifacts published on GitHub for replication.
The team analyzed nine services: ChatGPT, Claude, Grok, DeepSeek, Gemini, Perplexity, Microsoft Copilot, Mistral's Le Chat, and Meta AI. For each one they examined the web client in Chrome with full network capture, and the Android app using an instrumented phone that logs every outbound connection, permission access, and identifier read. They repeated the sessions across guest, free, and paid accounts, and across three consent states: ignore the cookie banner, reject all non-essential cookies, and accept all.
The conversations themselves were deliberately sensitive. The researchers chatted about health conditions and salary questions, because that is what real people actually ask chatbots, and they wanted to see what happens to that class of content.
Scale, for context: ChatGPT and Gemini each have more than 1 billion cumulative Android installs. These are not fringe apps.
The headline numbers
- Every single service contacted at least one third-party advertising or tracking service. No exceptions among the nine.
- 124 distinct third-party domains, from 44 organizations, were observed across the ecosystem. 34 of the 44 organizations are advertising or tracking services.
- 6 of 9 web clients and 3 of 8 Android apps leaked conversation-derived artifacts such as URLs, titles, prompts, or screenshots to third parties.
- Even after rejecting all non-essential cookies, 80.8 percent of third-party trackers stayed active. Reject-all is not the protection people assume it is.
- Free and paid tiers behaved nearly identically. Paying did not buy privacy here.
Google products had the broadest footprint, present in some form on all nine services. Sentry, Meta, Datadog, and Intercom followed. Nothing about this is unique to AI, by the way. These are the same pixels and SDKs that ship with every commercial web app. What is new is what they can see.
What exactly leaks: artifacts, not just clicks
This is the part that separates the study from routine privacy reporting. Traditional web tracking reveals where you browsed. Conversational AI generates artifacts that encode what you actually said.
- Conversation URLs. Five web clients disclosed conversation permalinks, or the identifiers needed to reconstruct them, to nine tracking organizations.
- AI-generated titles. Three web clients leaked conversation titles to nine third parties including Meta, TikTok, and DoubleClick. Titles matter because they are compressed summaries of your intent. The paper shows real examples: a prompt about early-stage Parkinson's symptoms became the title "Early-stage Parkinson's Symptoms." A mortgage question became "$85k NYC Salary: $280k-$350k Mortgage." That title is a health record or a financial profile, handed to an ad network alongside a tracking cookie.
- Prompts. On Grok's sharing pages, Meta Pixel received the user's latest prompt verbatim in its page-description field, and TikTok received it too.
- Screenshots. When a Grok conversation was shared, TikTok's pixel collected a screenshot of the most recent part of the conversation. The researchers captured the exact payload.
And these artifacts rarely travel alone. The paper documents Grok's web client sending the conversation URL and title to seven different trackers in a single session: DoubleClick, Google Ads, Google Search, Google Tag Manager, Meta, TikTok, and Twitter Analytics. Google Ads received a hashed user email alongside them. The Meta, TikTok, and Twitter cookies were synced across the flows, which means those companies can stitch the conversation into your existing profile on their platforms.
There is also a server-side layer that ad blockers cannot touch. Grok routes events through a server-side Google Tag Manager container that forwards conversation URLs and titles to Meta's and TikTok's server APIs directly. Claude's web client uses a first-party domain to load a configuration that forwards events server-to-server to eleven advertising platforms including Facebook, LinkedIn, TikTok, Reddit, and Google.
Anyone with the link could read your chats
The most alarming finding is about access control, not tracking. The researchers opened conversation permalinks in a fresh private browser session, logged out, and checked what was readable.
- Grok conversation URLs were publicly accessible by default on both free and paid tiers, with an opt-out available. Per the paper's disclosure timeline, this was still true on September 10, 2026.
- Perplexity made guest-tier conversations public with no opt-out at all. The authors note that Perplexity stopped sharing those URLs with trackers like Meta on April 3, 2026, possibly in response to a US class action, but the public-by-default accessibility remained.
- Every one of the nine services generates share links that work without authentication, and all nine share pages contained third-party trackers, giving those trackers full visibility into the shared conversation.
The team went one step further and planted canary tokens, unique URLs that phone home when accessed, inside conversations and uploaded documents. For Grok they recorded 70 activations from 70 different IP addresses across 48 autonomous systems in 14 countries, hours to days after the conversation ended. Two-thirds of those accesses came from the United States even though the conversations happened in the EU. Something was crawling shared conversations at scale. Only the initial access was observed for the other services, so the paper is careful to say the absence of further evidence there is not evidence of absence.
The uncomfortable competitive wrinkle
One detail that got less attention than it deserves: Google and Meta are embedded inside their competitors' products. Meta Pixel runs on Grok's web client. Google Analytics collects Gemini's conversation titles, and Google's ad stack sits inside ChatGPT, Claude, Copilot, and Perplexity. When a third party is embedded in a rival's chat product, the researchers note, the receiving company can potentially observe not just who you are but what the competing model told you. Your conversation is a competitive intelligence payload whether or not anyone has exploited it that way yet.
What rejecting cookies actually bought
The consent experiment produced the study's most practical lesson, and it is not a flattering one for the consent banners we all click.
- Accepting all cookies switched on additional trackers in Claude, Perplexity, and Grok: Meta, TikTok, Twitter Ads, DoubleClick, AppsFlyer.
- But ignoring the banner changed nothing compared to explicitly rejecting. Silent users got the reject-all tracking set, not something worse.
- Reject-all still left connections flowing in six services, including Google Ads calls from Perplexity, DeepSeek, Gemini, Copilot, ChatGPT, and Claude.
- Only ChatGPT's paid tier reliably remembered consent preferences across sessions.
Meanwhile, Android apps had no cookie banner equivalent at all. You accept the terms of service and the SDKs do what they do.
A privacy-focused browser that blocks known tracking domains at the network level, which is what the paper's mitigation section points to first, does more than any consent toggle on the page.
The five-minute settings checklist
The paper is careful to note its mitigations are partial and platform-dependent. Fair. Here is what I would actually do after reading it, ordered by impact.
- Block third-party tracking domains at the browser or network level. A browser with built-in ad and tracker blocking, or a network-level blocker, stops pixel and SDK calls regardless of what the consent banner says. This covers the majority of client-side leakage.
- Turn off or never use public chat sharing unless you truly need it. The paper showed share pages are the richest leak surface: prompts, titles, screenshots. If a chat contains anything sensitive, it should not have a public link.
- Assume conversation titles are public. Do not put identifying details in a first message you care about, because the auto-generated title will summarize them and travel with tracking payloads.
- Treat free-tier consumer chatbots as non-confidential. The paper's legal analysis under GDPR argues users' legitimate expectations hardly include Meta or Google reading their chats, but the technical reality on the ground today is that they might. Anything you would not paste into a public forum does not belong in a free consumer chatbot.
- On Android, reset the advertising ID and deny ad personalization. The study found resettable advertising IDs transmitted alongside persistent account IDs in several apps, which defeats the reset. It is still worth doing.
- Prefer incognito or temporary-chat modes where offered for sensitive questions. Several providers offer chats that are not saved to your history.
Why builders should care even if users do not
If you ship any LLM-powered product, a support chatbot, a wrapper app, an internal assistant with a web UI, the paper's scope section has a line you should read twice: these leakage vectors stem from embedding standard web and mobile analytics pipelines, so the concerns extend to custom chatbots and third-party AI wrappers built on the same stacks. Dropping a standard analytics SDK into your chat UI puts conversation titles and URLs on the wire to third parties whether you intended it or not. Compliance teams and enterprise customers are increasingly aware of exactly this class of finding, and the responsible-disclosure process here went all the way to European data protection authorities, with the Spanish DPA citing the work at an EDPB level.
The audit method itself is also worth stealing: the authors ran sessions with known-sensitive prompts and diffed the outbound traffic. You can do a version of this on your own product in an afternoon with browser dev tools and a HAR export, and you will probably find something you did not expect.
The takeaway
The paper's own conclusion is the right frame: these AI services challenge the perception of chat as a confidential exchange between you and a provider. The advertising web did not get replaced by AI. AI got absorbed into the advertising web, artifacts and all.
None of the nine services is clean. The differences are in degree. Grok's public-by-default permalinks and its screenshot-to-TikTok flow are the extreme end, while even the most restrained services in the study still contacted multiple third parties and leaked something under some consent conditions.
I write about AI systems, security, and the engineering around them every week. If this was useful, subscribe, it's free.
What is your setup? Have you checked what your favorite chatbot sends out when you press enter, or do you just accept the banner and move on? I am genuinely curious whether anyone has audited their own stack this way.
One thing I would do differently: before this paper, I would have said rejecting cookies was a reasonable baseline. It is not. The baseline now is blocking third-party domains at the network layer and assuming that anything typed into a free consumer chatbot may reach an ad platform with your identity attached.
Sources:
Top comments (0)