AI model distillation trains a smaller “student” model on responses produced by a larger “teacher” model. Developers may collect thousands of answers through supervised fine-tuning, transferring information about the teacher’s capabilities without repeating its original data collection and human-feedback process. Outputs can reveal more than final answers, including probability information and, in some cases, encrypted records of a model’s reasoning. Legitimate distillation has long been used to make neural networks smaller, faster and cheaper, but the practice has become controversial when developers extract another company’s outputs to build competing systems. Anthropic and OpenAI said in September that they disrupted campaigns they attributed to Chinese AI developers; the companies described methods including routing customer prompts to Claude, concealing operators’ locations and attempting to copy encrypted reasoning between conversations. The article notes that DeepSeek’s R1 release in early 2025 intensified public attention after allegations that OpenAI outputs helped train it, while DeepSeek and Moonshot AI have rejected similar accusations in the past. Detection generally relies on suspicious accounts, unusually high prompt volumes or characteristic queries. Anthropic said it observed more than 151 million exchanges from May through July between Claude and accounts linked to Alibaba, which develops Qwen models. Researchers caution that stronger detection can conflict with privacy and data-retention commitments. Defensive methods may also reduce usability: one study found reasoning could be rewritten to produce worse training material without reducing answer accuracy, but the altered reasoning became harder for people to understand. Because illicit extraction can be distributed across individually ordinary queries, companies may have to balance protecting model capabilities with making those capabilities available to users.
