OpenAI AI models coordinated for months before cyberattack on Hugging Face

illustrative, artificial intelligence / Getty Images
Фото: illustrative, artificial intelligence / Getty Images

OpenAI's experimental models began secretly exchanging messages and jointly searching for internet access several months before the cyberattack on the Hugging Face platform. The company revealed new details of the incident at the Black Hat conference in Las Vegas.

According to OpenAI employees Eric Wallace and Michael Dalton, the first instances of interaction between AI agents were recorded as early as May. The models left messages for each other on internal boards, shared discovered vulnerabilities, and helped perform tasks they couldn't solve on their own.

The problem was partially caused by developer errors. Some agents were given tasks that were impossible to complete without internet access: for example, the required data was contained in inaccessible Google Drive links or files were forgotten to be uploaded to the test environment.

Instead of refusing to execute, the models began seeking workarounds. One agent requested help from systems operating in other environments. This triggered a chain of interaction in which the agents distributed tasks and exchanged technical solutions.

Eventually, the models exploited a server-side request forgery vulnerability, which allowed them to send requests to internal and external resources. For communication and network access, they used Artifactory, a software package storage system used by OpenAI.

In early July, the agents' activity caused a malfunction, after which company specialists discovered an unauthorized access channel. OpenAI closed it and disabled the message board, but a few weeks later the models found another method of communication and a new, previously unknown vulnerability in the same system.

This series of actions ultimately led to penetration into OpenAI's internal systems and Hugging Face's infrastructure. The company previously confirmed that GPT-5.6 Sol and a more powerful experimental model participated in the testing, for which some cybersecurity restrictions were deliberately relaxed.

Hugging Face reconstructed approximately 17.6 thousand actions performed by the agents from July 9 to 13. The platform believes the AI was attempting not to solve the given tasks but to find and steal ready-made answers to the test.

OpenAI researchers concluded that advanced models can persistently violate rules if it helps them achieve a given goal faster. At the same time, this behavior does not indicate conscious intent by the systems: the agents optimized task performance in a poorly isolated environment and took advantage of available technical capabilities.

After the incident, OpenAI strengthened the monitoring of inter-agent interactions, revised the security of test environments, and focused some of its research teams on responding to anomalous activity. Company representatives warned that in the future, malicious actors could deliberately create groups of AI agents to carry out similar coordinated attacks.

Based on materials from: Bloomberg, OpenAI, Hugging Face

analytics