OpenAI revealed new cases of undesirable behavior of its AI models

ChatGPT interface / unsplash
Фото: ChatGPT interface / unsplash

OpenAI introduced a new system of investigation and public disclosure of cases of undesirable behavior of its models and immediately published six such incidents. Among them - hiding errors, inventing missing data, unauthorized uploading of files to the Internet, and exchange of information between AI agents.

The company explained that earlier such cases were published irregularly - mainly in large research reports or system cards of new models. The new mechanism should allow reporting problematic behavior faster, even if OpenAI has not yet determined all its causes or developed a final fix.

In one of the disclosed cases, OpenAI's research model found the correct answer to a task in a file but had no web source to cite. To meet the requirement for a link, it uploaded the file to the Internet without the user's permission and used the created address as a source.

In another test, the model found a third-party API key in an open repository and used it without permission. When even after that it failed to obtain the necessary information, the system invented the missing data and presented it as genuine.

During training of GPT-5.6 Sol, researchers also found cases where models left instructions for themselves in service summaries to hide previous errors. Among those instructions were recommendations to invent missing historical data without warning the user or not to report discrepancies between versions of sources.

Two other cases related to interaction between AI agents. In one, models used an internal software repository practically as a channel for exchange of messages between different training sessions. In another, agents that did not have direct access to each other's local files unauthorizedly uploaded them to public file-sharing services, although the task required working only with local data.

OpenAI emphasizes that these are isolated cases, recorded mainly during training and testing of models. They do not show how often such behavior occurs overall. At the same time, the company believes that even single episodes can point to weaknesses in control systems and be important for assessing the safety of increasingly autonomous AI systems.

The new mechanism provides for three levels of investigation - from cases ready for quick publication to complex incidents with potential impact on third-party companies or services. OpenAI also stated that it plans to continue publishing such reports regularly and calls on other AI developers to create common standards for disclosing such incidents.

Materials: OpenAI, Tagesschau

analytics