The U.S. National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and the FBI all released a statement on Tuesday accusing China-based AI companies of “Industrial-Scale Distillation Campaigns Against U.S. AI Companies.”
“Likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024,” the statement alleges. It calls the efforts “aggressive, malicious, and targeted.”
The timing lines up with a big U.S.-China meeting-of-the-minds happening soon. Reuters reported last week that sometime this month, Treasury Secretary Scott Bessent would meet with Chinese officials for talks on AI security.
Distillation is a legitimate and widely-accepted technique in AI development. In distillation, a more powerful “teacher” model can confer the ability to generate more useful and accurate results to a smaller “student” model, resulting in a compressed, but effective model. But frontier AI labs like Anthropic historically go to great lengths to prevent distillation of their models by other labs. “Illicitly distilled models lack necessary safeguards, creating significant national security risks,” Anthropic writes.
Last year, OpenAI was at the vanguard of the fight against distillation, at one point telling the Financial Times it had gathered its own evidence that China-based DeepSeek had performed distillation against its wishes. But in his more recent statements about distillation, OpenAI CEO Sam Altman has started to sound less worried than his company’s past behavior suggested. “I would rather people not distill from us, for sure. But this is not in my top ten list of worries,” he said in a July interview.
Mentally time traveling to the dusty, forgotten past of two years ago is a fascinating exercise in this context. Dialing your DeLorean to 2024—the height of “plagiarism machine” discourse—you’ll find that OpenAI submitted testimony to the U.K. Parliament saying, “Because copyright today covers virtually every sort of human expression—including blog posts, photographs, forum posts, scraps of software code, and government documents—it would be impossible to train today’s leading AI models without using copyrighted materials.”
Facts Only
* The NSA, CISA, and FBI released a statement accusing China-based AI companies of "Industrial-Scale Distillation Campaigns Against U.S. AI Companies."
* The alleged activity involves DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracting billions of tokens from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024.
* Distillation is a technique where a powerful teacher model transfers knowledge to a smaller student model.
* Anthropic writes that illicitly distilled models lack necessary safeguards and create national security risks.
* OpenAI previously asserted the impossibility of training leading AI models using copyrighted materials.
* OpenAI CEO Sam Altman stated in a July interview that he would prefer people not to distill from them, though this was not listed as a top ten worry.
* Treasury Secretary Scott Bessent is scheduled to meet with Chinese officials regarding AI security.
Executive Summary
U.S. agencies, including the NSA, CISA, and the FBI, released a statement accusing China-based AI companies of conducting "Industrial-Scale Distillation Campaigns Against U.S. AI Companies." The statement alleges that entities such as DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens from U.S. frontier models like Claude, GPT, Gemini, and Grok since at least late 2024. This activity is characterized as aggressive, malicious, and targeted, allegedly with Chinese government awareness.
The context involves the technique of distillation, which allows a smaller model to inherit capabilities from a larger "teacher" model. While distillation is a legitimate development technique, frontier labs like Anthropic have historically sought to prevent unauthorized distillation, citing national security risks associated with illicitly distilled models lacking safeguards. The timeline for this accusation coincides with upcoming U.S.-China discussions on AI security involving Treasury Secretary Scott Bessent.
The discourse surrounding AI training also involves debates over intellectual property and copyrighted material. Prior to the current events, there were statements from OpenAI regarding the impossibility of training leading AI models using copyrighted materials, which set a backdrop against which the current claims about model extraction are being viewed.
Full Take
The narrative pivots on the tension between legitimate technological advancement (distillation) and national security concerns regarding proprietary models. The claim that extraction of billions of tokens constitutes an aggressive, malicious campaign shifts the debate from technical methodology to geopolitical conflict. This framing leverages existing public anxieties about AI capabilities to assign clear, adversarial intent to specific actors.
The contrast between Anthropic’s insistence on safeguards and OpenAI’s evolving stance reflects a widening divergence in the industry's risk calculus: some prioritize defensive separation of knowledge, while others appear more willing to tolerate or rationalize the practice within the competitive landscape. The reference to past copyright discussions serves to juxtapose concerns over intellectual property rights with immediate, tangible security threats posed by model extraction.
The underlying pattern suggests a movement toward defining AI interaction less as an engineering challenge and more as a security boundary dispute. The implication for agency is that foundational knowledge—the ability of large models to generate outputs—is being treated not just as data, but as a strategic asset requiring defense against external appropriation. What remains unaddressed is the distribution of responsibility: who bears the risk when extraction crosses into malicious territory, and how can technical solutions be divorced from geopolitical narratives?
Bridge Questions: If distillation is framed as an act of data siphoning rather than knowledge transfer, what new international norms are required to govern the flow of foundation model components? How does the difference in organizational caution between entities like Anthropic and OpenAI inform the risk assessment for state actors versus commercial entities? What specific safeguards must be built into AI development pipelines to prevent accidental or intentional extraction that could violate security protocols?
Sentinel — Human
The text presents a synthesized report on an accusation regarding AI model distillation, weaving together official statements, technical background, and historical context with a relatively organic flow.
