Your Chat Title Is an Ad Impression: Nine AI Assistants Measured End to End

The threat model most people carry for a chatbot is “the provider can read my conversation.” A measurement paper that reached a wide audience this week argues the real model is broader and duller: the provider hands a compact, semantically rich summary of your conversation to the same advertising infrastructure that already tracks the rest of your browsing, and in some cases hands over a URL that lets the recipient go read the whole thing.

“Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents” is authored by Guilherme Oliveira, Miguel Sanchez, Juan Manuel De Santa Olalla Gómez, Roi S. Serna, Tautvydas Jackevicius, Jorge Garcia-Herrero, Aniketh Girish, Guillermo Suarez-Tangil and Narseo Vallina-Rodriguez — predominantly IMDEA Networks, formatted for Proceedings on Privacy Enhancing Technologies. It is the full write-up behind the LeakyLM disclosure the authors released in May. The measurements were performed manually in Spain during May 2026 against nine services: ChatGPT, Claude, Grok, DeepSeek, Perplexity, Gemini, Microsoft Copilot, Mistral Le Chat and Meta AI — all nine web clients, and the eight that ship an Android app.

The base rate: everyone, on both platforms

Across web and Android the authors observed 124 distinct third-party domains attributable to 44 organisations, of which 34 are advertising and tracking services (ATSes). Every single evaluated service contacts at least one. Eleven ATSes appear on both client types; fifteen are web-only, including Google Tag Manager, TikTok and the consent platform OneTrust; eight are mobile-only, including Braze.

One structural detail deserves more attention than it usually gets: 71% of all distinct endpoints contacted by the mobile clients originate from WebViews rather than native bytecode — 98% in Grok, 91% in Perplexity, 71% in ChatGPT. The “app” is frequently a browser with web tracking inside it, running under the app's permissions, without the app store's consent surface. If your mental model treats mobile as the more locked-down client, this inverts it.

What actually leaks is the summary, not the transcript

The novel finding is not that trackers are present. It is what they receive. Conversational AI generates artifacts that encode the subject matter: the auto-generated conversation title, the permalink, the conversation ID, sometimes the raw prompt. Overall 6 of 9 web clients and 3 of 8 Android clients disclose conversation artifacts to third parties — to 11 ATSes on web and 2 on mobile.

  • Conversation URLs: 5 web clients disclose the permalink or its global conversation identifier to 9 ATSes. Three of those five do it by default, without any consent interaction.
  • Conversation titles: 3 of 9 web clients leak the AI-generated title to 9 third parties including Meta, TikTok and DoubleClick. 8 of those 9 disclosures occur only after the user accepts non-essential cookies — which is the system working roughly as designed, and also the point at which a user who clicked “accept” to dismiss a banner starts broadcasting what they are asking about.
  • ChatGPT and Claude send the globally unique conversation ID to Datadog as a standalone parameter. The authors note the ID permits reconstruction of the public URL.

The title is the sharpest artifact here. It is a model-generated one-line precis of the conversation's topic, and the authors deliberately probed with health-related prompts precisely because that category is a GDPR special category. A title is smaller than a transcript and far more legible to an ad-targeting pipeline.

Grok is the outlier, and it is not close

Two Grok behaviours stand out. First, Grok embeds a server-side Google Tag Manager container that transmits the conversation URL and chat topic title server-to-server to Meta's Conversions API and TikTok's Events API — carrying both Meta's _fbp and TikTok's _ttp cookies in the same payload, which is textbook identity bridging across two ad platforms. Because the forwarding happens server-side it is invisible to the browser and unblockable by content blockers.

Second, when a Grok conversation is shared, the authors observed TikTok collecting screenshots of the conversation and both Meta and TikTok collecting the user's most recent prompt. That is verbatim conversation content, not metadata. Grok's permalinks are also the most permissive of the nine: publicly accessible by default on free and premium tiers, with opt-out rather than opt-in.

Claude's server-side path is worth naming too: Segment Analytics proxied through the first-party domain a-cdn.anthropic.com, loading a Conversion API configuration that forwards user events server-to-server to eleven trackers, again evading ad blockers. This only activates after accepting non-essential cookies — rejecting them prevented Meta Pixel, Datadog telemetry and the eleven-way forwarding from firing.

The canary tokens are the part that turns exposure into access

Leaking a URL only matters if someone dereferences it. The authors tested that directly, embedding canary URLs in conversations and in uploaded .docx and PDF files, using a self-hosted canary on a generic domain because some models recognise and decline known canary-token domains.

For Grok they recorded 70 separate canary activations from 70 distinct IP addresses spanning 48 autonomous systems across 14 countries and 4 continents, hours to days after submission. The conversations were conducted in the EU; 65.7% of the retrievals came from machines in the United States. For DeepSeek, Copilot, Mistral and Claude, activations were generally one-off at submission time from AWS and GCP ranges. Perplexity repeatedly fetched canary URLs even when the prompt explicitly instructed it not to, via its Perplexity-User crawler on US AWS infrastructure.

The authors are appropriately careful that absence of observed access is not evidence of absence. But for Grok the positive result is unambiguous: shared conversation resources remain subject to distributed automated retrieval long after the interaction ends.

Privacy controls barely move the number

This is the finding operators and policy people should carry forward. Even when users reject non-essential cookies, 80.8% of third-party trackers remain active. Perplexity, DeepSeek, Gemini, Copilot, ChatGPT and Claude all still connect to Google Ads under reject-all. Separately, 77.3% of web-identifier transmissions to third parties occur only after acceptance — so consent does shift a large share of identifier flow, while leaving the baseline connection set largely intact.

Subscription tier is close to irrelevant: free and premium accounts contact near-identical third parties. And the authors make a pointed observation about who the remaining control is available to — on web, only ChatGPT's paid tier supports persistent consent preferences, while roughly 95% of ChatGPT users are on free tiers and OpenAI has announced advertising for free and low-cost users. The strongest control is furthest from the most exposed population.

On identifiers, the mobile side is the more durable problem. Grok transmits email addresses and hashed email addresses to TikTok and Twitter Analytics; Perplexity sends email to RevenueCat and hashed emails to Singular. Three mobile clients send the Android Advertising ID to third parties — Adjust (Copilot), AppsFlyer (Grok), Singular (Perplexity) — and in the Grok and Perplexity cases the AAID travels alongside persistent user or installation IDs, which defeats the whole point of a resettable identifier: reset it and the install-level ID re-binds the new one. Claude sends geolocation coordinates and the Android ID to Sift Science. DeepSeek brings two Chinese fingerprinting and risk-scoring services, Fengkong Cloud on mobile and ShuMei on web, that are absent from the usual tracker databases.

The regulatory track is already moving

The disclosure timeline in the paper is unusually complete, and it is the part that distinguishes this from a one-off measurement:

  • 23 March 2026 — tracker activity first noticed in Perplexity and Grok traffic.
  • 13 April 2026 — findings notified to EU and UK data protection authorities.
  • 17 April 2026 — xAI notified via its vulnerability disclosure address about Grok's conversation exposure. This is the one issue the authors classify as an exploitable security flaw rather than an intentional product decision.
  • 4 May 2026 — partial findings published as LeakyLM.
  • 27 May 2026 — the Spanish DPA, AEPD, forwarded a note on the study to the European Data Protection Board and requested it be circulated to every European DPA and raised at the plenary meeting of 8–9 June 2026.
  • 15 August 2026 — OpenAI updated ChatGPT's privacy policy to explicitly mention third-party trackers. The authors state plainly that they cannot confirm a causal link to their disclosure.
  • 10 September 2026 — Grok still uses publicly accessible permalinks. No response received from xAI.

Perplexity did change behaviour: it stopped sharing conversation URLs with third-party trackers on 3 April 2026, which the paper links to a US class action filed on 31 March 2026. Guest-tier Perplexity conversations remain public.

The legal analysis argues the practices sit awkwardly against Article 5(3) of the ePrivacy Directive and the GDPR's transparency and legal-basis requirements, noting that providers describe these flows only through generic terms — “user content” (OpenAI), “conversations” (Anthropic), “service interaction info” (Perplexity), “user info” (xAI) — where CJEU case law requires per-operation disclosure of data, purpose and legal basis. The authors are explicit that this is issue-spotting, not a compliance determination.

What to do with this

  • Treat consumer assistant conversations as third-party-visible by default. Not the transcript necessarily, but the topic. If a conversation title would be damaging as an ad-targeting signal, it does not belong in a consumer tier.
  • Never share a conversation permalink containing anything sensitive. Sharing is the mode where verbatim prompts and screenshots reached third parties, and where the canary evidence shows subsequent retrieval.
  • Reject non-essential cookies anyway — it demonstrably suppressed the highest-value flows on Claude and Grok — but do not mistake it for containment, because four in five trackers survive it.
  • If you build on these platforms, the same pipelines are in your wrapper. The authors are clear the leakage vector is standard web and mobile analytics embedded in a conversational interface, which generalises to support chatbots and third-party AI wrappers. Putting a Meta Pixel on a page that renders model output is the whole bug.
  • Enterprise tiers were explicitly out of scope. Providers market them with distinct commitments that this study did not validate. Do not assume the findings transfer either way.

The limitations are stated honestly and matter: nine services, one country, one month, black-box analysis, single-session interactions, no desktop or voice clients, no Gemini Android trace. The presence of a third party does not by itself prove an advertising purpose — some of these SDKs do crash reporting and payments. The authors call their results a point-in-time lower bound, which is the right frame.

It also lands in a pattern this site has tracked for a while: the AI layer keeps inheriting the failure modes of the layer beneath it rather than inventing new ones. Grok shipping a build that uploaded whole git repositories, Chrome exposing a silently installed on-device model to any page, and now conversation permalinks flowing into ad-tech identity graphs are all the same shape — a new, semantically dense data type dropped into old plumbing that was never designed to be careful with it.

Sources: