The AI Model Distillation Paradox
Summary
The article argues that the U.S. government applies inconsistent standards to AI training and model distillation. In a Justice Department brief supporting AI companies in copyright litigation, the government said ingesting copyrighted text to train models can be fair use and warned that licensing requirements could weaken U.S. innovation. Eight days later, the NSA, CISA, and FBI accused six Chinese laboratories, including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.ai, of industrial-scale campaigns that extracted billions of tokens from American frontier models such as Claude, GPT, Gemini, and Grok. Distillation trains a smaller student model to imitate a larger teacher model, reducing data, compute, latency, and hosting costs; the author argues that it is a specialized form of training whose main difference is that the source is another model rather than human-created material. U.S. companies also use distillation, including examples such as Alpaca, Orca, and Grok, while their terms of service generally prohibit customers from training rival models on their outputs. The legal status remains unsettled: reverse engineering is generally lawful, copyright usually does not protect AI-generated outputs, and trade-secret claims are uncertain, although alleged conduct may breach contracts. The article says American labs have not sued despite reported large-scale campaigns because litigation could create rules that also constrain their own practices. It highlights the contradiction that U.S. labs have themselves harvested books, videos, and other material without permission, while Anthropic has both accused Alibaba of distillation and agreed to a $1.5 billion settlement with authors over training books without permission. The author also notes that distilled Chinese open-weight models can return to U.S. development pipelines, creating a boomerang effect. Arguments about market substitution, access-control evasion, free-riding, lost safeguards, and national security may describe real risks, but they do not by themselves establish a consistent legal distinction. The article concludes that the U.S. has proposed no clear rule defining the legal status of model outputs, leaving a policy conflict between encouraging transformative innovation and restricting capability transfer.