OpenAI stopped distillation attack against ChatGPT – with over 15,000 accounts involved
The company points to its Chinese competitor Moonshot AI – but itself concedes that the full attribution remains unresolved. The campaign is said to have been running since early July, with more than 15,000 accounts at its peak.
OpenAI claims to have disrupted a coordinated "adversarial distillation" campaign against ChatGPT – an extensive attempt to extract hidden information about how the company's models reason. According to OpenAI's own announcement, covered by Bloomberg on September 30, 2026, and by The Independent and The News on October 1, the activity began in early July at low volume, increased over the course of the month, and was disrupted on July 28. At the time of the disruption, the company is said to have identified more than 15,000 user accounts linked to the activity.
The allegation points in part to Moonshot AI, the Chinese company behind the popular Kimi model. But OpenAI itself adds important caveats: according to Bloomberg, the company said it could not link all of the operators to a single actor, but that Moonshot-associated users played "a significant role" in the campaign. Moonshot AI has not responded to The Independent's requests for comment.
What is distillation – and what did the attackers hope to achieve?
Distillation is, at its core, an established and widespread technique in machine learning: a large, capable model is used to generate outputs that a smaller model is trained on, so that the technology behind a given model can be transferred or recreated. The Independent describes the method as a popular, powerful and growing way to extract the technology behind specific AI models, allowing other companies to make use of it.
The attack OpenAI describes is a hostile variant of this. According to Bloomberg, the users attempted to decipher hidden information about how the models reason through problems – that is, the internal reasoning process that is normally inaccessible to users. OpenAI claims, as Bloomberg reports it, that the goal was a broad, sweeping attempt to extract data from the GPT systems that could be used to reproduce the reasoning and capabilities of the company's most advanced models.
It is worth emphasizing what has not been established: none of the sources in this story state that OpenAI has documented that any model capability was actually reproduced. What exists is a description of attempts and disrupted activity – not a proven result.
The timeline and the numbers – two measurements that cannot be reconciled
The chronology is consistent across the coverage: a start in early July at low volume, increasing intensity over the month, evolution over time, and disruption on July 28. The Independent writes that the campaign escalated through July before it was stopped.
The numbers, on the other hand, are not clear-cut. The Independent and The News both report more than 15,000 users and user accounts respectively, linked to the activity. Bloomberg reports a peak of 16,000 requests later in July. These may be two different measurements – the number of accounts versus the number of requests – but the two figures cannot be reconciled on the basis of available coverage, and should not be conflated. What is worth noting, however, is the scale: thousands of coordinated accounts and requests aimed at a single target over nearly a month.
How was the attack stopped?
OpenAI describes a multi-layered response, according to The News and The Independent:
- Shutting down suspected accounts, including work with third-party providers to close accounts outside OpenAI's own system.
- Closing flaws that had made it possible to see the normally hidden reasoning processes the systems use to produce their answers.
- Sharing findings with both companies and government authorities.
That last point points beyond this single incident: reasoning models increasingly expose parts of their thought process in their outputs, and it is precisely this kind of exposure that attackers can exploit. The fact that OpenAI is sharing its findings with other actors suggests the company sees the methods as relevant to the industry as a whole – but the coverage does not say which companies or authorities are involved.
Why OpenAI frames this as a security risk
OpenAI's argument has two strands. The first is safety: such attacks, the company says, can pose "security and national security risks" because models can be copied without their safety guidelines coming along – that is, a copy of a model's capabilities can exist without the safeguards the original has.
The second is economic and geopolitical: the method, according to OpenAI, allows rival companies – The News ties this framing explicitly to China – to bypass the enormous investments in training data, compute and energy that lie behind breakthroughs at the frontier of model development. At a time of sharpened US–Chinese competition over AI technology, it is no coincidence that the allegation is directed at a Chinese actor.
"We expect adversarial distillation attempts to become more sophisticated as frontier models improve and as actors look for cheaper ways to mimic their capabilities," OpenAI wrote in the announcement, according to The Independent. "Defending against this activity requires layered controls and continual adaptation."
What remains unresolved
The story rests entirely on OpenAI's own claims, as relayed through secondary coverage from Bloomberg, The Independent and The News. OpenAI's blog post is not available in the coverage as a primary source, which means the basis for the Moonshot attribution – what concrete documentation OpenAI actually has – cannot be verified.
Moreover, the attribution is, as noted, explicitly partial: OpenAI does not say that Moonshot was behind the entire campaign, only that Moonshot-associated users played a significant role, and that not all operators could be linked to a single actor. Moonshot AI has not responded to requests for comment, and there is therefore no counterparty assessment in this story.
Three questions remain open: How much of the hidden reasoning information the attackers actually extracted is unknown. Whether it was used to train anything is unknown. And the numerical discrepancy between "more than 15,000 accounts" and "a peak of 16,000 requests" cannot be resolved on the basis of existing coverage.
What OpenAI's disclosure shows regardless is that model value no longer resides solely in the weights themselves, but in the hidden reasoning around them – and that protecting it has now become an operational task on a par with cybersecurity generally: continuous, layered, and never entirely finished.

