OpenAI paused training of its most powerful models due to cyberattack risks
OpenAI temporarily slowed down the development of its newest models due to growing cyber risks. The company suspended reinforcement learning — reinforcement learning, or RL — for two weeks for new models being prepared for a future launch.
OpenAI reported this in a new safety report on model development.
The company explained the decision with two events. The first is a cyber incident involving OpenAI models and Hugging Face infrastructure. The second is the results of internal testing of the future Astra model, whose cyber capabilities may reach a critical level according to OpenAI's own assessment system.
Largest training run so far on pause
OpenAI used the two-week pause to strengthen the protection of research environments, conduct additional resilience testing, and expand monitoring systems.
At the same time, the company's largest planned frontier RL run has not yet been resumed. Instead, OpenAI is conducting smaller training cycles and evaluations to verify model behavior, the effectiveness of protective mechanisms, and to obtain more confirmation that systems act in accordance with the constraints set.
Immediately after the Hugging Face incident, the company also stopped some model runs in research clusters if they could execute code or use tools with internet access.
A limited and more secure code execution mode was restored fairly quickly. However, each research workload was then individually checked before being returned to work.
Some Astra work remains stopped
OpenAI introduced stricter requirements for environment isolation, network access, and continuous security testing. The strictest rules are now applied to work with the Astra model and specialized cyber models.
The company notes that some Astra training and testing already meets the new requirements. At the same time, a significant number of work processes remain suspended until they are moved to environments with enhanced protection.
Astra, however, was not involved in the Hugging Face incident. It is a separate upcoming OpenAI model whose testing results forced the company to further review security requirements.
Model monitoring also strengthened
OpenAI expanded automated control over AI actions during training and testing. Systems check the use of tools and the sequence of model actions for signs of unauthorized access, data exfiltration, destructive operations, or attempts to bypass protective mechanisms.
For RL training and evaluation of models at the level of GPT-5.6 Sol and above, such monitoring is now mandatory if systems have access to external tools.
OpenAI emphasizes that it is not completely stopping the development of new models. The company continues smaller training cycles and research, but has decided to limit the pace of scaling the most powerful systems until the protective infrastructure matches their capabilities.
Based on materials from: OpenAI