NSA, FBI and CISA accuse six Chinese AI firms of industrial-scale distillation of American models
The US intelligence and security agencies NSA, FBI and CISA have accused six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI — of extracting billions of tokens from American frontier models.

NSA, FBI and CISA accuse six Chinese AI firms of industrial-scale distillation of American models
The US intelligence and security agencies NSA, FBI and CISA have accused six Chinese AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI — of extracting billions of tokens from American frontier models. The accusation came roughly ten days before Treasury Secretary Scott Bessent and China's Vice Premier He Lifeng meet in New York for negotiations where AI is explicitly on the agenda. But while the authorities present distillation as the core of China's AI progress, experts disagree over whether the claim actually holds up.
What the accusation involves
According to a joint advisory from NSA, the FBI and the Cybersecurity and Infrastructure Security Agency (CISA), designated AA26-251A, the six companies are conducting "aggressive, malicious and targeted distillation activities at industrial scale." It is worth being clear from the outset: the advisory's primary text is not available in this reporting — every concrete detail is known through secondary coverage, particularly from TechTimes and CNBC.
As relayed by TechTimes, the advisory accuses:
- DeepSeek of using outputs from four versions of Claude, two versions of Gemini, five versions of ChatGPT and Grok 4 to improve capabilities such as agentic reasoning, code generation and professional writing.
- Moonshot AI of distilling 18 different American models into its own systems Kimi K2 and Kimi K3.
- Alibaba, MiniMax, StepFun and Z.AI — without model-specific details of the same kind given in the coverage.
Technically, the advisory describes extraction of "billions of tokens" from American frontier models. The level of detail in the accusations — down to the number of model versions — is striking, but that is precisely why the absence of the primary text is a problem for readers trying to assess credibility: we do not know what the evidence base is, how it was obtained, or what documentation the agencies actually hold.
What distillation is — and why the controls don't capture it
Knowledge distillation is a well-established technique in machine learning: a smaller or newer model is trained to imitate the outputs of a larger, established model. It does not take the teacher model's weights or code — it learns from what the model produces. A teacher generating answers to millions of tasks serves as the training signal for a student that never sees the textbook.
This is what places distillation outside the existing export-control regime. The American chip controls target hardware: advanced GPUs and the equipment needed to train frontier models. But if an actor can acquire capability through an API key — paying for access to ChatGPT or Claude and using the answers as training data — neither imported chips nor access to the original model are required. The ease that makes APIs a business model also makes them a leak.
This is the core of the enforcement gap the accusation exposes: America's main AI policy tool toward China controls hardware, while the capability actually being transferred can pass through a paid service. Advisory AA26-251A can be read as an attempt to give API-based extraction the same moral and political weight as chip smuggling — a framing shift with consequences for how such activities can be enforced in trade agreements.
The timing: right before the negotiations
Bessent, US Trade Representative Jamieson Greer and He Lifeng are meeting in New York, according to a statement from Greer's office relayed by Reuters. The meeting comes ahead of the summit between President Donald Trump and China's President Xi Jinping on September 24. Sources tell Reuters that AI governance is among the items on the agenda, alongside rare earths.
Bessent has confirmed to Axios, as reported by TechTimes and corroborated by Reuters, that the talks will explicitly cover "both open and closed weight models." That is a fairly technical distinction to bring to top-level diplomacy: open models with freely available weights (which Chinese actors largely export) versus closed, API-based models (dominated by American labs). That the distinction is now on the negotiating table means the distillation question — which concerns how capability moves between precisely these two systems — is no longer just a security matter but a negotiation topic.
China's Ministry of Commerce has dismissed the accusations as "groundless" and warned of countermeasures if Washington uses the distillation issue to restrict Chinese AI companies, according to TechTimes. The statement sets the frame: if the US tries to turn distillation into an enforceable trade clause, China will treat it as an escalation.
The contested core: How much does distillation actually matter?
This is where the story departs from most security accusations: even within the American and Western debate, the central premise of the claim is disputed.
CISA has claimed that Chinese AI companies are "conducting systematic extraction of proprietary features and capabilities from American AI companies' models through industrial-scale distillation campaigns that form the core—not merely a supplement—of their AI development strategy," according to the quote CNBC relays. Anthropic went further in a report published in September, claiming that companies including Alibaba, Moonshot and DeepSeek — among China's leading AI labs — attempted "illicit distillation" to train their own models. Note that this is a company claim, not independently verified, and that Anthropic is itself among the potentially affected parties.
The pushback comes from prominent people with direct industry insight. Aidan Gomez of Cohere called Chinese AI models "world class" and said the lead American labs hold is "evaporating very quickly," according to CNBC. The point is indirect but important: if Chinese models are world class and the gap is closing quickly, that suggests they have capabilities distillation alone cannot explain.
Sriram Krishnan, a former senior policy advisor for AI at the White House, went further. Services like ChatGPT and Claude, he told CNBC's "Squawk Box," "came out of distilling human content." It is a reminder that distillation is a core technique across the entire field — Americans distill the web, and every model builder distills competitors' outputs where the terms of service allow it. Krishnan suggested that distillation's actual contribution to Chinese progress is unclear.
The contradiction is real and should be presented as unresolved: the security agencies and Anthropic present distillation as the very engine of Chinese AI development; two weighty technical voices present it as a common, often legitimate practice that hardly explains the entire gap it is supposed to explain. Neither side has, in this coverage, presented quantitative evidence that settles the matter.
The contradiction in American practice
The story also has an internal tension on the American side. Reuters reported that a US government website uses an AI model from a Chinese company to let users search proposed federal regulations — at the same time as the FBI accuses the model's developer, Alibaba, of "malicious" copying of Anthropic's technology. The American executive branch is thus using a model from a company simultaneously described as a threat to American technology.
This kind of contradiction is not necessarily evidence against the accusations — security assessments and practical operational decisions can be separated — but it weakens the coherence of the message. It becomes harder to present distillation as existential theft when the accused model is also approved as a tool in the American administration.
What is unverified — and what to watch
Readers should weigh several uncertainties:
- The primary source is missing. Advisory AA26-251A is not available, and every specific claim — six companies, model versions, "billions of tokens" — is known through TechTimes, CNBC and Reuters. TechTimes in turn cites Axios and other underlying reports not available here.
- The date is not established. The sources do not establish the advisory's exact publication date, only that it came roughly ten days before the New York meeting.
- The companies' responses are missing. None of the accused companies has in this reporting commented directly on the accusations — only China's Ministry of Commerce has responded on the industry's behalf.
- The evidence behind the specific model claims — that DeepSeek used exactly four Claude versions and five ChatGPT versions, that Moonshot distilled exactly 18 models — has not been made public.
The three things that will settle the matter: the outcome of the New York negotiations and the September 24 summit, and whether concrete proposals for API- or distillation-related measures emerge there; whether the advisory's primary text and evidence are made public; and whether US enforcement moves from chips toward API usage — which would change both American terms of service and Chinese negotiating cards.
For now, the safest thing to say is that the accusation is serious and precisely formulated, but unproven in public — and that its most important function so far may not be law enforcement, but negotiating positioning just before the table is set.