Anthropic has accused Alibaba of running the largest illicit model distillation campaign it has ever measured, part of a wave of extraction attacks from seven China-based AI labs that the company says it detected and disrupted since February 2026. The allegations appear in Anthropic's September 2026 threat intelligence report, published Thursday, which dedicates an entire section to industrial-scale attempts to clone Claude's capabilities.
The disclosure is the most detailed public accounting yet of a practice that has quietly become one of the AI industry's most contentious problems. For anyone tracking breaking AI news, the report reads as an escalation in the ongoing battle between US frontier labs and Chinese competitors over who controls the most capable models — and how those capabilities are obtained.
What Anthropic Says Alibaba Did
According to the report, operators affiliated with Alibaba's Qwen, or Tongyi, Lab injected a fixed prompt into every request that forced Claude to write out its full chain-of-thought reasoning inside inline text tags before answering. Those reasoning transcripts were then saved and converted into supervised fine-tuning data, which Anthropic says was used to help train Qwen 3.5, 3.6 and 3.7.
The scale is striking. Anthropic says the campaign targeted the chain-of-thought transcripts of Opus 4.6 and 4.7, peaked at nearly 3 million exchanges per day launched from more than 3,500 fraudulent accounts, and totaled over 151 million observed exchanges between May and July 2026. The company states the operation ran through two pools of nearly 5,000 fraudulent accounts that used residential proxies, disposable email addresses and virtual card payments to hide their tracks. When Anthropic banned the first pool, the traffic simply shifted to the second.
Beyond distillation, the report alleges Alibaba also used Claude to advance its own AI research and development, including building reinforcement learning environments and model architecture research. Anthropic says some accounts were found to be funneling requests from DeepSeek and Xiaomi as well, suggesting the same proxy networks serve multiple organizations.
Moonshot Served Claude Instead of Kimi
The allegations against Moonshot AI, the maker of the Kimi model family, are perhaps the most unusual. Anthropic reports that Moonshot silently forwarded customer requests to Claude instead of processing them with its own models — then displayed Claude's responses to users who believed they were talking to Kimi.
In one instance, Anthropic counted almost 300,000 customer requests relayed to Anthropic over just ten days, the vast majority routed to Opus. The operation ran through a proxy network of 5,380 fraudulent accounts, most appearing to be located in Singapore and Japan. Between May and July 2026, Anthropic attributes more than 23 million exchanges to the campaign.
Moonshot also built what Anthropic calls a cross-session replay attack: saving the encrypted "thinking signature" Claude returns instead of raw reasoning, then starting a new session and prompting Claude to convert that signature back into the full reasoning trace. Anthropic says DeepSeek used the same technique, and that new defenses against it are being deployed.
The report flags a serious privacy dimension: user queries Moonshot rerouted included sensitive information its customers never consented to sharing with a third party. Examples include a user likely affiliated with China's People's Liberation Army analyzing CCTV surveillance data on a targeted individual from hundreds of cameras in Chengdu, and an engineer at a major Chinese state-owned enterprise whose internal code and live credentials passed through Claude without his knowledge.
DeepSeek's Relay Operation
Anthropic says DeepSeek deployed similar tactics, silently relaying user requests to Claude Opus and harvesting reasoning traces via the same cross-session replay method. The company reportedly identified targets by checking inbound requests for strings associated with third-party coding harnesses like Claude Code, the Claude Agent SDK and OpenCode, then selected tagged users' requests for relay.
The scale attributed to DeepSeek: over 12.1 million exchanges observed across just 14 days in July 2026. The relayed data exposed what users surely assumed was private. Anthropic's examples include an employee of a Chinese technology company analyzing internal documentation for a flagship AI program, an IT operator working with data from a Russian government agency associated with its Ministry of Defense whose requests exposed live database credentials, and engineers at a municipal Public Security Bureau in China building a tool that tracks individuals' movements against police records by national ID number.
Zhipu, Xiaomi and the Wider Pattern
The report also attributes a chain-of-thought extraction campaign to Zhipu AI, which brands itself as Z.ai outside China. Anthropic says Zhipu ran its pipeline against Opus 4.8 by rotating through 273 fraudulent accounts over ten days in June, counting 770,609 exchanges through its CoT-extraction cleaner and over 3 million total exchanges in the same period.
DeepSeek, Xiaomi and Moonshot are separately accused of feeding conversations between their own users and models into Claude as training data. Anthropic says those sessions contained names, email addresses, company data and other sensitive information from hundreds of end users in at least a dozen languages — much of it relayed through third-party model routing services popular with users in the United States and Europe — and that the practices are likely inconsistent with privacy laws and the labs' own terms of service.
Why Distillation Is So Hard to Stop
Distillation itself is a legitimate and widely used training method: a smaller "student" model learns from a larger "teacher" model's outputs. What Anthropic defines as illicit is the industrial-scale, covert version — fake accounts built on stolen credit cards, credentials and API keys, plus purchased user transcripts from proxy services that save conversations without users' knowledge.
The extraction attempts documented in the report range from blunt to creative. One lab ran a 12,000-request experiment, each request using a different technique, to find which ones would extract Claude's reasoning. Another directed the model to "translate" its previous working memory into katakana-only Japanese to slip past anti-distillation filters.
OpenAI has publicly flagged distillation attacks since early 2025, and Google published a threat tracker on adversarial distillation earlier this year. Anthropic's contribution is scale and specificity: named labs, named models, and hard numbers.
Anthropic says all seven labs' attacks targeted its generally available models, and that it has observed no attempts against its restricted Mythos 5 or Mythos Preview systems, which are not accessible to the public. It also notes that models distilled from frontier systems inherit none of their safety guardrails — and that its own research suggests distillation can lift dangerous capabilities in biological and cyber domains even when the harvested exchanges contain little about those subjects.
The company says every campaign described was disrupted, that safeguards have been strengthened in response, and that intelligence was shared with authorities and industry partners where appropriate. Whether naming names changes the economics of cloning frontier models — or merely documents an arms race already in full swing — will depend on what happens next.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →