OpenAI reported a mass attempt to extract data from its models

OpenAI reported a mass attempt to extract data from its models
Photo: OpenAI / illustrative

OpenAI said it detected and stopped a large-scale coordinated campaign in which participants tried to obtain hidden reasoning of its AI models. The company linked the main group of such activity to people associated with Chinese Moonshot AI - the developer of a family of models and the Kimi chatbot.

According to OpenAI, the first signs of activity appeared on July 1. Initially, the number of requests was small, but on July 24-25, the company recorded a sharp surge - about 16 thousand requests from more than 4 thousand users, which matched the pattern of extracting protected reasoning.

After expanding the investigation, OpenAI discovered related activity with similar request patterns in a cluster that included more than 15 thousand users. According to the company, by July 28, this activity was completely stopped.

OpenAI characterizes such actions as hostile model distillation. Distillation itself is a common AI development method in which the output of a more powerful model is used to train another. However, in this case, according to OpenAI, it was unauthorized systematic acquisition of data that could help reproduce or improve the capabilities of a competing model.

Campaign participants, as OpenAI claims, tried to obtain not just the models' ordinary responses, but their protected internal reasoning. One method involved moving encrypted reasoning data from one conversation to another and asking the model to decrypt and represent them in an accessible form.

At the same time, the company emphasized that the attackers did not break encryption, did not gain access to databases, and did not penetrate stored user conversations. It was about manipulating interaction with models through a large number of specially crafted requests.

Bloomberg Law notes that OpenAI could not determine whether all campaign participants acted on behalf of a single entity. At the same time, the company stated that the main cluster of activity is linked to people associated with Moonshot AI.

Moonshot AI is one of the largest Chinese developers of generative AI. The company develops the Kimi line, and its current flagship model Kimi K3 is aimed at complex tasks, programming, and logical reasoning.

After the incident, OpenAI closed the mechanism that allowed reuse of encrypted reasoning fragments, added new checks for potential leakage of such data, and began blocking associated accounts. The company also shared information about the discovered scheme with other AI developers through the Frontier Model Forum and with government structures.

OpenAI expects that attempts at unauthorized model distillation will become more sophisticated as advanced AI systems develop. The company plans to strengthen technical protection, systems for detecting coordinated activity, and information exchange about such threats with other market participants.