Bessent says sanctions and Entity List designations are on the table
The National Security Agency and the Federal Bureau of Investigation were among the agencies that on Tuesday accused six Chinese companies of training their AI systems by drawing on American models. The agencies said the campaigns amounted to “aggressive, malicious, and targeted distillation activities at an industrial scale.”
The accusation comes as administrations on both sides prepare for a meeting between President Trump and Chinese leader Xi Jinping at the end of the month in Washington, where AI issues could be on the agenda.
Distillation works by letting a new model learn the behavior of an existing one. A developer such as Anthropic or OpenAI trains a model by feeding it large amounts of text; the model builds a weighted network of links, then needs an additional training step before it can answer in natural ways. In this polish step, engineers can prompt an AI “teacher model” for answers and explanations, then use those question-answer-reasoning sets to fine-tune a “student model.” That fine-tuning process is called distillation.
The agencies said the six companies had engaged in millions of exchanges with frontier U.S. models, including Claude, ChatGPT, Gemini, and Grok. Anthropic said in February that Chinese operators were using “fraudulent accounts and proxy services to access Claude at scale while evading detection.” In July, an Anthropic executive said distillation had helped China narrow the gap with the U.S. from 12 to 18 months down to roughly six to nine months.
An executive at Beijing-based Moonshot AI told Chinese media in July that the Kimi K3 release’s “breakthrough performance” was based on fundamental innovations, not distillation or copying. A Chinese foreign ministry official said the same month that foreign countries were hyping the concept of distillation from malicious motives.
Some U.S. researchers, including OpenAI executive Dean Ball, say distillation might have helped Chinese companies earlier in the AI race but isn’t the main reason for their recent advances.
By itself, distillation probably falls short of theft, according to some legal specialists. Model output is unlikely to be copyright-protected the way a book or a movie is, and even if it were, distillers are not copying output outright, merely using it in the way an aspiring writer learns by reading books. Some distillation is accepted, such as Apple’s distillation of its own model to run on phones, and distillation that works with open-weight, accessible models is widely accepted. The tension arises when a Chinese company distills closed U.S. models, including Claude.
Anthropic’s terms of service bar Chinese companies from using Claude, and other firms have similar restrictions. Lawyers advising U.S. companies say such firms could sue Chinese rivals for breaching the terms of service, arguing that the low-cost copied models damage their sales. But evidence is difficult to obtain if the conduct was hidden via middlemen overseas, and a lawsuit could be slow in a quickly-moving industry.
Treasury Secretary Scott Bessent, writing on X in July, said: “When PRC [People’s Republic of China] firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.”
The security agencies called for closer cooperation between companies and the government to defend against large-scale distillation, but did not discuss specific sanctions on China. If the blacklist threat is carried out, U.S. and foreign users might lose access to the Chinese models, a measure that would hurt Chinese developers.
A July open letter from companies including Nvidia and Microsoft cautioned against “sweeping restrictions on techniques that play an important role in AI innovation” and asserted that unlawful distillation “should be addressed through targeted legal and commercial frameworks.”