190 Million Stolen Exchanges: The Anthropic Distillation Accusation — and the Beijing Probe of Its Own Labs

Anthropic's fourth threat-intelligence report, Detecting and countering misuse of AI: September 2026 — a 154-page document published September 10 covering activity disrupted between December 2025 and August 2026 — levels an accusation with no precedent at this scale: seven China-based labs ran industrial-scale illicit distillation campaigns against Claude, generating a combined roughly 190 million exchanges whose outputs were allegedly fed into their own training pipelines. Anthropic names Alibaba, Moonshot AI, DeepSeek, Zhipu, MiniMax, Xiaomi, and SenseTime. This site covered the report's cyber-operations findings on September 15; the distillation half has since broken open into a regulatory confrontation neither side can walk back.

The turn came the week of September 22. Citing The Information, The Next Web reported that China's internet regulator, the CAC, is investigating DeepSeek and Moonshot over Anthropic's allegations — specifically the claim that the labs secretly routed millions of user exchanges through Claude to train competing models. Decrypt's September 23 account frames the trigger plainly: Anthropic's report is what forced Beijing's hand. A regulator probing its own national champions on a foreign competitor's evidence is an unusual event, and it signals that large-scale distillation has crossed from terms-of-service dispute into state-level data-governance territory. Note the evidentiary chain here: Anthropic's telemetry is provider-side and self-interested, the labs have not publicly substantiated their training-data lineage, and the CAC probe is confirmed only through press reporting — but the probe itself is now a fact enterprises must plan around regardless of how the underlying claims resolve.

Distillation is a supply-chain problem now

The security relevance is not who copied whom. It is that a large fraction of the open-weight and API-accessible ecosystem may carry unknown training lineage. A model distilled at industrial scale from a frontier system inherits capabilities — including agentic coding and tool-use behaviors that reporting ties directly to this case — without inheriting the safety evaluations, refusal tuning, or audit trail of either the source or the claimed training process. Enterprises adopting such models absorb unpriced risk: unclear license and terms-of-service taint, untested failure modes, and potential leverage for the regulator that claims jurisdiction over the data flows. Google's threat tracker has flagged distillation as an adversarial-use category since early in the year; Anthropic's report is the first to attach nine-figure exchange counts and named labs to it.

The DeepSeek angle compounds the concern. Separate research notes on the report allege selected customer requests were relayed to Claude Opus with the exchanges collected for training — the SaaS-wrapper-as-collection-vector. And Unit 42 has already documented DeepSeek-linked tooling in autonomous attack campaigns. Whether or not any single allegation survives scrutiny, the pattern is that the model supply chain — weights, APIs, wrappers, resellers — is now an active collection surface, and procurement can no longer treat a model card as provenance.

What to do

  • Demand training-data lineage in procurement, and treat its absence as a finding. Ask vendors which base models and data pipelines produced the system you are buying, including distillation sources. A model card that cannot answer is not a bargain — it is an unscoped dependency.
  • Enforce your own API boundaries against extraction. Rate-limit, anomaly-detect, and contractually prohibit bulk output harvesting on any AI service you expose. If seven labs can pull 190 million exchanges from a frontier provider before disruption, your unmonitored internal deployment is not going to notice a smaller operation.
  • Track the CAC probe as regulatory risk, not distant politics. If you build on DeepSeek, Moonshot, or other named-lab models, map what a forced retraining, weight withdrawal, or export restriction would do to your roadmap — and keep a second-sourced alternative warm.
  • Watch for wrapper and reseller collection. Audit the AI-adjacent vendors with access to your prompts: resellers, wrappers, and "cheap Claude" intermediaries. Anthropic's report documents attackers pulling production API keys out of wrapper services — your key in someone else's pipeline is their training data.

The uncomfortable arithmetic: roughly 190 million exchanges, seven labs, one 154-page report — and now a regulator investigating its own champions. Model capabilities have become extractable at industrial scale through nothing more exotic than the public API, and the only parties with full visibility into the flows are the providers being harvested and the states claiming jurisdiction. Everyone else is buying lineage they cannot verify.

Sources: