When you chat with an artificial intelligence assistant, what you read on the screen is only part of the process. Behind it lies an internal reasoning that AI companies guard jealously, because it reveals how the model reaches its conclusions. OpenAI has confirmed that someone attempted to misappropriate that reasoning on a large scale, and that it managed to stop it after months of surveillance.
The matter is relevant now because on September 30, 2026, OpenAI published a detailed statement about what happened, in which it places the origin of the problem in the first week of July of that year. The company speaks of a coordinated operation that involved more than 15,000 users and which it attributes, at its core, to individuals linked to Moonshot AI, the Chinese company behind the Kimi model.
What the campaign organizers were looking for
OpenAI describes what happened as a case of adversarial distillation. The term refers to the systematic and unauthorized use of responses or reasoning from another model to train, reproduce, or improve another artificial intelligence system.
The specific goal was protected reasoning, that is, the internal log that the model generates while processing a task before providing a final answer. According to the company, that content may include information that is never shown to the user, and which would be valuable to anyone wanting to replicate the model's capabilities without starting from scratch.
OpenAI insists on an important nuance: there was no intrusion in the classic sense. The operators did not break the company's encryption, did not access any database, and did not manage to enter directly into stored conversations of other users.
Timeline: from July to the peaks on the 24th and 25th
According to OpenAI's account, the activity started on July 1 with a low, almost undetectable volume. The situation changed drastically on July 24 and 25, when 16,000 requests with a clearly identifiable extraction pattern were recorded, coming from more than 4,000 different users.
Subsequent investigation broadened the focus and detected related prompt activity in a much larger group of over 15,000 users. OpenAI maintains that it managed to completely interrupt the operation on July 28.
The company clarifies that these figures reflect extraction attempts, not necessarily cases in which the extraction was successfully completed. It also acknowledges that the attack pattern changed over time, which, according to its own analysis, requires maintaining layered defenses that constantly adapt to new techniques.
The method: encrypt, move, and ask to decrypt in another conversation
The technique described by OpenAI is striking for its apparent simplicity. Operators copied fragments of encrypted reasoning generated in one conversation and moved them to a different conversation, where they asked the model to decrypt and transcribe them.
In this way, they managed to have content intended to remain hidden reproduced in a format readable to the requester. The operation was repeated in a coordinated and large-scale manner, in what OpenAI qualifies as a direct breach of its terms of service.
The company adds a relevant detail: independent security researchers had previously communicated to it, through responsible disclosure channels, vulnerabilities related to the exchange of information between models and conversation compaction processes. OpenAI confirmed that those attack vectors were real and acknowledges that this external work helped it better understand the type of threat and accelerate its mitigation measures.
Attribution to Moonshot AI: certainties and unclear areas
OpenAI is cautious when it comes to assigning responsibility. The company itself acknowledges that it does not have absolute certainty as to whether all the users involved were acting under the coordination of a single actor.
Despite that caution, the statement does point out that a central group of the activity is associated with individuals linked to Moonshot AI, the company responsible for the Kimi model. OpenAI does not publicly detail what specific evidence or indications led it to that conclusion.
The company also emphasizes that this type of manipulation is not exclusive to its own models. For this reason, it decided to share the information gathered with other companies in the sector through the Frontier Model Forum, a collaboration platform between major artificial intelligence developers.
Why OpenAI talks about national security
OpenAI's central argument is that the extracted reasoning could be used to train a different model without maintaining the safeguards that the company applies to the visible responses of the original.
On a large scale, OpenAI maintains, this type of distillation could also accelerate the transfer of advanced artificial intelligence capabilities without the recipient having to assume the same investment in security required to develop a model from scratch. The company notes that this concern increases as AI models gain capabilities in dual-use fields, that is, applicable to both civilian and potentially sensitive purposes.
It should be kept in mind that this assessment comes from the affected company itself, which also competes directly in the market with the actors it points to.
OpenAI’s response: blocked accounts and new technical barriers
OpenAI combined several lines of action. At the account level, it blocked or restricted access to profiles considered fraudulent, reinforced controls applied to the registration of new users, and increased surveillance on networks of related accounts.
At the technical level, the company closed a specific path that allowed someone who already possessed another user's encrypted reasoning to reproduce it and recover its content. It also introduced additional checks to detect and block streaming-format outputs that could expose fragments of protected reasoning.
When suspicious activity was channeled through third-party services, OpenAI worked with those providers to identify and deactivate the involved accounts. Finally, it shared its findings with both the Frontier Model Forum and government information-sharing channels on security.
What those who use or integrate OpenAI models should review
Several practical recommendations for those who develop products on its models can be extracted from the company's own statement.
The first is to check whether one's own system stores or forwards reasoning artifacts in a portable or reproducible way, something OpenAI identifies as a potential risk.
The second affects those who deploy models through a cloud partner: it is advisable to verify that such deployment has the same protections as the direct service, as the company itself admits that these third-party hosted environments need equivalent safeguards.
The third recommendation is aimed at applications that chain different AI tools: OpenAI warns that attacks that take advantage of the output of those tools require protections capable of examining more than just the text visible to the end user.
Finally, the company advises monitoring anomalous usage patterns, since the detected campaign relied precisely on fraudulent accounts and volumes of requests outside the ordinary.
What the user can expect from now on
The statement makes it clear that the investigation has not been closed. OpenAI continues to analyze and mitigate these types of attacks, and anticipates that adversarial distillation attempts will become increasingly sophisticated as artificial intelligence models themselves improve.
For those who use these systems, whether as a private user or as a developer through the API, the practical consequence points to more controls on account registration, greater surveillance of unusual usage patterns, and possible additional restrictions on how reasoning data is managed. Those who maintain active integrations with OpenAI models would do well to review their configurations and the current terms of service.

