Do AI Chatbots Share Your Conversation Data With Trackers? The Nine-Service Study, Explained
A peer-reviewed study found conversation-derived artifacts reaching third parties from six of nine web chatbots. Here is what it proved—and what it did not.
Do AI Chatbots Share Your Conversation Data With Trackers? The Nine-Service Study, Explained
Short answer: A peer-reviewed study of ChatGPT, Claude, Grok, DeepSeek, Perplexity, Gemini, Microsoft Copilot, Mistral Le Chat, and Meta AI found third-party advertising or tracking services in every product tested. In May 2026 measurements, conversation-derived artifacts reached third parties from 6 of 9 web clients and 3 of 8 Android apps. Those artifacts ranged from conversation IDs and URLs to generated titles—and, in one Grok sharing flow, the latest prompt and a screenshot.
That does not mean every chatbot sent every full conversation to advertisers. The products, platforms, recipients, consent states, and exposed fields differed. The most alarming verbatim-content result belonged to a specific Grok sharing flow. Several other findings involved analytics or support vendors receiving identifiers, URLs, or AI-generated titles rather than raw prompt text.
Research cutoff: October 8, 2026.
The Study's Main Finding in One Minute
The accepted PoPETs 2027 paper, “Prompt like a Butterfly, Sting like a Tracker”, examined nine popular AI chat services on the web and eight corresponding Android apps. The researchers ran controlled accounts and prompts from Spain, captured client traffic, inspected third-party requests, varied cookie choices and account tiers, and tested whether conversation links opened without authentication.
The headline numbers are:
- All 9 services integrated at least one service the researchers classified as advertising or tracking infrastructure.
- 6 of 9 web clients and 3 of 8 Android apps disclosed some form of conversation-derived artifact to a third party.
- 5 web clients disclosed conversation URLs or reconstructable conversation IDs to 9 third-party services.
- 3 web clients disclosed AI-generated conversation titles to 9 third parties.
- 80.8% of measured third-party trackers remained active after non-essential cookies were rejected.
- The researchers found no clear, general reduction in third-party data collection between free and premium tiers.
The paper was deposited by IMDEA Networks in September after earlier preliminary disclosure and has been accepted for Privacy Enhancing Technologies Symposium 2027. Fresh October coverage and community discussion have renewed attention around the results.
The correct takeaway is not “all AI chats are publicly leaked.” It is that the familiar web and mobile tracking stack can receive artifacts that are unusually revealing when the page is an AI conversation.
What Is Confirmed
The researchers tested nine recognizable AI services
The web study included:
- ChatGPT
- Claude
- Grok
- DeepSeek
- Perplexity
- Gemini
- Microsoft Copilot
- Mistral Le Chat
- Meta AI
Eight Android apps were also tested; Gemini's mobile traffic could not be extracted successfully, so the mobile denominator is eight rather than nine.
The experiments took place in Spain during May 2026. Researchers used fresh browser profiles, controlled test accounts, an instrumented Android device, and predefined prompts involving privacy-sensitive health scenarios. They inspected network requests, request bodies, cookies, browser storage, device identifiers, account identifiers, common encodings, and known hashed values.
This matters because the conclusions come from observed traffic, not a privacy-policy keyword search. It also matters because the findings are a point-in-time measurement from one region, not an eternal description of every current build.
“Conversation data” covers several very different artifacts
The study's umbrella category includes:
- a global conversation ID;
- a private or public conversation URL;
- a share-link ID or URL;
- an automatically generated conversation title;
- a prompt;
- a screenshot preview;
- identifiers that can associate the artifact with an account, browser, device, or session.
A generated title may summarize a sensitive topic even when it is not a verbatim prompt. A conversation URL may reveal no readable content when access controls are strong—or may expose the whole conversation when anyone with the link can open it. A conversation ID may look less revealing, but the paper says some IDs allowed reconstruction of a public URL.
Those distinctions are why “chatbots send chats to advertisers” is too broad. The actual privacy risk depends on what was sent, to whom, under which interaction, and whether the referenced conversation was accessible.
The ordinary web flows differed by provider
The paper's detailed table reports these notable normal-conversation web flows:
- ChatGPT sent a conversation identifier and related account, device, and session identifiers to Datadog; in the tested web flow, the table also records a chat URL.
- Claude sent conversation identifiers to Datadog and, in some configurations, a conversation URL plus account information to Intercom.
- Gemini sent an AI-generated conversation title to Google Analytics.
- Grok sent a conversation ID, URL, and generated title to seven advertising or analytics endpoints after non-essential cookies were accepted.
- Mistral Le Chat sent a conversation ID, URL, title, and account-linked fields to Intercom after cookie acceptance.
- Perplexity sent a conversation ID and URL with account or session identifiers to Datadog.
The paper did not report the same artifact pattern for DeepSeek, Copilot, or Meta AI on the web. That does not mean those products had no third-party components; it means the study did not observe the same conversation-artifact disclosure in those tested flows.
Several recipients in the table—such as Datadog and Intercom—provide monitoring, support, or operational tooling. Others—such as Meta Pixel, TikTok Pixel, DoubleClick, and Google Ads—fit the advertising label more directly. Treating every recipient as an ad network would erase an important distinction.
The strongest raw-content finding involved Grok sharing
The most direct content exposure appeared after a Grok conversation was shared. The paper reports that:
- Meta and TikTok received the latest user prompt through share-page metadata.
- TikTok received a screenshot preview of the most recent part of the conversation.
- The share URL, generated title, prompt, and tracker cookies could travel together.
That is materially different from a conversation identifier reaching an error-monitoring service. It is also narrower than the claim that all ordinary Grok prompts, or all prompts in every tested chatbot, were transmitted to advertisers.
The researchers say Grok's normal conversation page also sent the conversation URL and generated title to seven tracking or analytics services when non-essential cookies were accepted. Their appendix shows the same conversation identifier propagating to Google Ads, DoubleClick, Google Tag Manager, Meta, TikTok, Google Search, and X Analytics.
Rejecting cookies helped, but it did not end third-party activity
Consent made a meaningful difference in some web clients. The paper says rejecting non-essential cookies in Claude prevented Meta Pixel, Datadog telemetry, and server-side forwarding to eleven advertising platforms from activating in the tested configuration. Grok's seven-endpoint normal-conversation flow was also observed only after cookie acceptance.
But rejection was not a universal off switch. Across the measured configurations, 80.8% of identified third-party trackers remained active after non-essential cookies were rejected. Four of nine free-tier web services still had third-party trackers collecting data in that condition.
That 80.8% figure should not be misquoted as “80.8% still received full conversations.” It measures tracker activity, not the percentage receiving raw chat content. Some active services received operational or identifier data rather than prompts.
Paying did not consistently remove the tracking surface
The researchers found free and premium tiers interacted with largely the same third parties. They observed one mobile exception—Intercom and Sentry appeared in Claude's free tier but not the premium tier—but no general paid-tier privacy transformation.
Paying for a chatbot may change limits, models, ads, storage commitments, or enterprise terms. In this study's consumer configurations, it did not consistently remove third-party infrastructure.
Enterprise and government tiers were explicitly outside the study. You should not apply these consumer-client findings to a contractual enterprise deployment without checking that product's data-processing terms and technical controls.
Sharing a chat usually creates a bearer link
All nine services supported a user-triggered sharing flow that produced a conversation page accessible without logging in. That is usually how a share link is intended to work: possession of the link grants access.
The privacy problem begins when the page also loads third-party code, when the URL is sent elsewhere, or when users mistake an unguessable link for a private, access-controlled document. The researchers found nine third parties embedded across conversation-sharing pages, giving those services visibility into the rendered pages.
The paper separately reports that Grok's normal conversation permalinks were readable by anyone with the URL by default, with an opt-out, and that Perplexity guest conversations were public by default in the tested build. The authors say Perplexity stopped sending those guest URLs to third-party trackers on April 3, 2026, before the final paper.
Regulators were notified
The researchers say they notified relevant European and UK data-protection authorities on April 13 and xAI on April 17. Spain's data-protection authority later published a notice asking European authorities to examine whether some AI systems let third parties access conversations. The AEPD notice reflects regulatory interest, not a final violation finding.
The paper says Grok still used publicly accessible permalinks on September 10 and that the researchers had received no official response to their xAI disclosure by that date.
What Is Still Unclear
Which measured behaviors still exist today?
The traffic capture occurred in May. AI products change rapidly, sometimes several times a week. A provider may have removed a request, changed a vendor, added a consent gate, altered a share page, or moved processing server-side since the measurement.
The accepted paper preserves valuable evidence of what the researchers observed, but it is not a live scanner. We found no public, same-harness October rerun covering all nine services and every tested account state. Current users should treat the platform-by-platform details as verified historical measurements that need fresh replication—not as a promise that nothing changed.
Why did each third party receive the data?
The paper deliberately classifies advertising, analytics, attribution, monitoring, fraud, and support services under a broad advertising-and-tracking umbrella. Its limitations section says the presence of a third party does not prove the data was used for advertising under every condition.
Datadog may receive telemetry for reliability. Intercom may support customer service. A fraud service may evaluate abusive sign-ins. Those purposes still create a data-boundary question, especially when a URL or account identifier is involved, but they are not identical to building an advertising audience.
The study observed the transfer. It could not inspect every recipient's internal use, retention, onward sharing, access controls, or contract.
Did a human read anyone's private conversation?
The study used researcher-controlled accounts and did not access other people's chats. It demonstrated that certain artifacts were transmitted and that shared or permissively accessible pages could be opened.
The authors embedded canary URLs in their own conversations to detect later automated access. They recorded 70 Grok canary activations from 70 IP addresses across 14 countries, sometimes hours or days later. Perplexity also fetched test URLs repeatedly through its crawler, including when a prompt asked it not to.
Those callbacks show that downstream infrastructure accessed test resources. They do not identify a human reader, prove that every conversation was fetched, or establish that private user data was sold.
How representative was one run per configuration?
The researchers say pilot tests showed deterministic tracking behavior, so each configuration was analyzed once. That is a defensible engineering choice for a large matrix, but it limits estimates of intermittent behavior, experiments, regional rollouts, and account-specific flags.
The study also excluded iOS, native desktop apps, voice-first interfaces, enterprise tiers, government tiers, and long-lived memory behavior. Gemini mobile traffic could not be extracted. Results may differ by geography and over time.
Is the legal analysis final?
No. The authors analyze the observed flows under the GDPR and ePrivacy Directive and argue that transparency, legal basis, consent, and special-category-data questions deserve scrutiny. The AEPD escalated the issue for broader European review.
But neither the paper nor the AEPD notice is a final enforcement decision. A binding compliance finding would require facts the researchers could not see, including contracts, purposes, internal retention, security controls, and provider responses.
What You Should Do Before Sharing Sensitive Data With an AI Chatbot
1. Separate content from identity
Remove names, email addresses, patient numbers, account IDs, employer details, case numbers, and unique dates when they are not necessary. A scrubbed medical question can still be sensitive, but it is harder to attach to a person than the same question paired with direct identifiers.
2. Treat chat titles as data
An AI-generated title can reveal the subject even when the prompt is not transmitted. “Possible pregnancy complication,” “bankruptcy options,” or “employee misconduct investigation” may be sensitive on its own.
Before pasting confidential material, ask whether a generated summary, page title, URL, or telemetry event could expose the topic.
3. Avoid public share links for confidential work
A share link is designed for access by whoever has it. Do not use one for medical records, legal strategy, credentials, proprietary code, personnel issues, or regulated data unless the product provides—and you verify—appropriate access controls.
If you already created a link, revoke it in the originating service. Deleting the chat may not revoke a separately created share page unless the product says it does.
4. Reject non-essential cookies, but understand the limit
Rejecting cookies reduced several measured flows and is still worthwhile. It did not eliminate all third-party activity in the study, and a browser blocker cannot see provider-to-provider server-side forwarding.
Cookie choices, model-training controls, chat-history controls, memory settings, share-link access, and provider retention are separate layers. One toggle rarely controls all of them.
5. Check the exact plan and surface
Consumer web, mobile, enterprise, API, and government products can have different terms and architectures. A vendor's enterprise “no training” commitment does not automatically describe its consumer web telemetry. A mobile app result does not automatically describe its API.
Document which product, plan, region, device, and settings you are approving for sensitive use.
6. Use the narrowest tool that fits the task
If you only need help rewriting, outlining, comparing public facts, or brainstorming, do not hand an autonomous system your email, cloud drive, browser session, or public-sharing authority. Fewer connected systems create fewer paths for sensitive context to travel.
For high-risk material, use an approved enterprise environment with reviewed contracts and controls, or keep the sensitive fields out of the request entirely.
Where OpenVeil Fits—and Where It Does Not
OpenVeil is a privacy-focused hosted AI workspace for adults. In normal OpenVeil conversations, reopenable chat history is kept in the browser rather than as a normal server-side chat-history account record. Documented OpenVeil product content is not used to train foundation models.
That creates a different normal-history boundary for people who want a conversational workspace without a conventional server-side account archive of their chats.
The boundaries matter:
- OpenVeil is hosted, not fully offline.
- Active requests still require processing by OpenVeil and necessary providers.
- OpenVeil is not anonymous and does not promise zero logs.
- OpenVeil is not a browser tracker blocker, cookie manager, ad blocker, VPN, endpoint-security tool, or public-link access-control product.
- OpenVeil cannot remove data already sent to ChatGPT, Claude, Grok, Gemini, or another service.
- OpenVeil does not make it safe to paste secrets, credentials, regulated records, or another person's private information without authorization.
The natural OpenVeil use case is narrower: a person needs AI help with ordinary analysis, drafting, or exploration and prefers normal reopenable chat history to stay browser-local instead of becoming a normal server-side chat-history account record.
Frequently Asked Questions
Did the study prove ChatGPT sends full prompts to advertisers?
No. The tested ChatGPT flows sent conversation identifiers and related account, device, session, and URL fields to Datadog. The paper did not report ordinary ChatGPT prompt text being sent to an advertising network.
Did the study prove Claude shares conversations with Meta or TikTok?
Not in the broad way that sentence implies. The paper observed Claude third-party infrastructure and reported that rejecting non-essential cookies prevented Meta Pixel, Datadog telemetry, and server-side forwarding to eleven advertising platforms in the tested setup. Its conversation-artifact table lists identifiers and a URL going to Datadog or Intercom, not Claude prompt text going to Meta or TikTok.
Which chatbot result was most serious?
The clearest verbatim-content exposure involved Grok's sharing flow: Meta and TikTok received the latest prompt, and TikTok received a screenshot preview. Grok also had permissive normal-conversation permalink access in the tested build.
Does rejecting cookies stop AI chatbot tracking?
It can reduce it, but the study did not find that it stopped all third-party activity. The reported 80.8% figure refers to measured trackers that remained active after rejection, not the percentage still receiving full conversation content.
Does paying for a chatbot make it private?
Not automatically. The consumer free and premium configurations in this study contacted largely similar third parties. Enterprise plans may have different contracts and controls, but they were not tested.
Are AI chat share links private?
Usually they are bearer links: anyone who has the URL can open the shared page. That is not the same as a private document restricted to named recipients. Revoke the link when it is no longer needed and do not use it for confidential content.
Is OpenVeil fully offline or free from provider processing?
No. OpenVeil is hosted, and active requests require processing by OpenVeil and necessary providers. Its relevant distinction is the documented browser-local boundary for normal reopenable chat history and the commitment not to use documented product content for foundation-model training.
Bottom Line
The study supports a serious but precise conclusion: mainstream AI chat interfaces have inherited the tracking and analytics machinery of ordinary websites and apps, and that machinery can receive unusually revealing conversation artifacts.
It does not show that every chatbot sends every full chat to advertisers. The measured exposure ranged from IDs and URLs to titles, prompts, and screenshots. The strongest raw-content finding was tied to Grok's sharing flow; other products showed different, narrower transfers.
The practical response is to treat titles, URLs, share pages, identifiers, and telemetry as part of the privacy boundary—not just the text box. Reject unnecessary cookies, avoid public share links for sensitive work, remove identifiers, verify the exact plan and surface, and use the narrowest tool that can do the job.
Sources
- Oliveira et al., “Prompt like a Butterfly, Sting like a Tracker” (accepted PoPETs 2027 paper)
- IMDEA Networks: Your conversations with AI may not be as private as you think
- IMDEA Networks publication record
- AEPD notice on third-party access to AI conversations
- Open research artifacts
- Current October 7 coverage summarizing the accepted paper