OpenAI published a security post on September 30, 2026 describing how it identified and disrupted a coordinated campaign to extract protected reasoning from its models — and attributed a core cluster of the activity to individuals associated with Moonshot AI, the Chinese lab behind the Kimi model family.
The disclosure lands at a sensitive moment for the industry, where the value of a frontier model increasingly lies not just in its answers but in how it reasons. For readers tracking the latest AI developments, the case offers an unusually detailed look at how model IP is attacked at scale — and how labs are responding.
What "adversarial distillation" means here
OpenAI describes the activity as consistent with adversarial distillation, which it defines as the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model. "Protected reasoning" is the model's internal record for working through a task — information that is normally withheld from the final answer a user sees. Extracting it, the company argues, can help others reproduce capabilities they did not build.
Importantly, OpenAI says the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester — a coordinated, high-volume abuse of the product itself rather than a traditional hack.
One technique OpenAI documented involved copying encrypted reasoning traces from one conversation and then asking a model in a separate conversation to decrypt and transcribe the hidden reasoning content. The company stressed the manipulation is not a vulnerability unique to its own models, and said it shared details with industry partners through the Frontier Model Forum to strengthen collective defenses.
The timeline: 16,000 requests in two days
According to the post, the campaign began on July 1, 2026 at relatively low volume. It escalated sharply on July 24 and 25, when OpenAI observed high-volume spikes of 16,000 requests using a relevant extraction pattern, coming from more than 4,000 users. A footnote in the post cautions that these figures describe attempted extractions — not necessarily successful ones.
Further investigation uncovered related prompt-pattern activity across a cluster of more than 15,000 users. OpenAI says it fully disrupted the campaign by July 28, 2026, and that the activity continued to evolve over time, reinforcing its view that adversarial distillation requires layered, adaptive defenses rather than a one-time fix.
Attribution points to Moonshot AI
OpenAI is careful with its attribution. It says it is unclear whether all operators observed during the period came from a single actor, but that it attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
The claim does not arrive in a vacuum. On September 8, 2026, the National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the FBI released a joint cybersecurity advisory stating that Moonshot AI has conducted a widespread distillation campaign against U.S. frontier AI companies since at least mid-2025. According to the advisory, Moonshot extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train Kimi-K2, using U.S. models to distill capabilities in supervised fine-tuning, reinforcement learning, software engineering, and mathematics.
The advisory goes further, stating that — likely with Chinese government awareness — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI have collectively extracted billions of tokens across millions of exchanges from U.S. frontier models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024. The agencies describe a logistics pipeline of native APIs, remote cloud providers, and third-party aggregators, plus a gray market of API proxies known as "transfer stations" used to bypass geographic restrictions and terms of service. Cost savings reportedly came from bulk procurement of premium subscriptions shared across developer teams.
Why reasoning traces are worth stealing
The security concern goes beyond competitive advantage. OpenAI argues that extracted reasoning could be used to train another model without preserving the safeguards applied to the original model's user-facing outputs — meaning distillation at scale can transfer advanced capabilities without the corresponding investment in safety. The company says those risks grow as models gain abilities in dual-use domains.
Independent researchers have been sounding the same alarm. In a paper submitted to arXiv on August 10, 2026, a group of researchers including Alexander Panfilov, Ilia Shumailov, and Maksym Andriushchenko reported that leading providers do not fully conceal their models' reasoning traces, and demonstrated related cross-model and conversation-compaction vulnerabilities through responsible disclosure. OpenAI says it investigated those findings, confirmed the attack paths were real, and used the work to accelerate its mitigations.
What comes next
OpenAI says it investigated the campaign's scope and potential impact before publishing, deployed its own mitigations, and shared its findings with researchers and industry partners for feedback. Additional mitigation and investigation work is continuing, and the company frames adversarial distillation as a broader security challenge for the entire frontier-lab ecosystem rather than a single incident.
The episode also sharpens an ongoing policy debate. If U.S. agencies are formally attributing distillation campaigns to Chinese labs, expect tighter controls on API access, more aggressive rate-limiting and identity checks, and continued pressure on the aggregator and proxy markets that make large-scale abuse easier to hide. For AI companies on both sides of the Pacific, the reasoning layer is now a contested asset — and protecting it has become a security discipline of its own.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →