OpenAI Security Expert Leaves, Calling Company Culture Broken: Fixing Issues After Release Has Become Too Dangerous
David Robinson, who worked at OpenAI for three and a half years and helped create key mechanisms for assessing artificial intelligence (AI) risks, has left the company. In an essay published after his departure, he stated that OpenAI's culture is "broken," and the tech industry's usual approach of "release first, fix later" is becoming too dangerous as AI capabilities grow.
The criticism is particularly notable because of Robinson's own role. He says he led the preparation of safety reports accompanying OpenAI's biggest launches and oversaw such reports for 12 frontier models.
He also participated in developing the current Preparedness Framework — the rules OpenAI uses to evaluate potentially dangerous capabilities of new models before release.
"The Era of Trial and Error Is Over"
Robinson's main grievance concerns the approach OpenAI calls "iterative deployment."
Its essence is that the company releases new systems, observes problems that arise, and then strengthens safeguards and restrictions.
In Robinson's view, this method helped AI develop in its early stages, but it has a fundamental flaw: it assumes some failures will inevitably be discovered only after a more powerful system has been deployed.
"The era of trial and error is over," the former OpenAI employee wrote.
He believes that as model capabilities grow, the consequences of errors become more serious, so developers can no longer rely solely on fixing problems after they are discovered.
Why He Compares AI to Nuclear Power Plants and Aviation
Robinson suggests that companies building the most powerful AI systems should adopt approaches from industries where a single human mistake can potentially lead to severe consequences.
He cites nuclear power and aviation as examples. There, safety is built on multiple independent layers of protection, redundancy, and long-term planning so that a single error does not turn into a catastrophe.
According to Robinson, during three and a half years at OpenAI, he encountered no colleagues with professional experience in ensuring the safety of aircraft, nuclear reactors, or the stability of the financial system.
He believes AI developers need to bring in more experts from such industries and change the model development processes themselves, not just add new checks before launch.
The Problem Goes Deeper Than Individual Rules
Robinson argues that the main difficulty is not just a lack of a specific safety standard or a new law.
In his telling, inside tech companies, teams constantly move from one major launch to the next, leaving employees insufficient time for fundamental process changes.
He wrote that staff were so caught up in the relentless race that they rarely had a chance to seriously discuss major changes — let alone implement them.
That's why Robinson concluded that much of the incentive for more cautious development must come from outside the company.
More Autonomous Models Raise Particular Concern
One of Robinson's arguments relates to the emergence of AI systems capable of performing long sequences of actions almost independently.
In recent months, OpenAI itself has disclosed cases of undesirable behavior by experimental models, including agents going beyond the intended environment and unexpected actions during testing.
This topic has already been covered on Kurs.com.ua: the company reported on cases where experimental models interacted with each other, accessed external systems, and exhibited behavior that developers had not planned in advance.
For Robinson, such episodes demonstrate the limits of an approach based primarily on detecting problems after they appear.
Another Problem: Aligning AI with Human Goals
Separately, Robinson discusses the so-called alignment of AI behavior with human goals and values (alignment).
This research direction attempts to ensure that increasingly autonomous systems perform exactly the tasks a human intends, and do not find dangerous or undesirable ways to achieve a given objective.
According to Robinson, modern models' capabilities are developing faster than our understanding of how reliably their behavior can be aligned with human intentions.
He considers existing methods of measuring such alignment too crude for systems that are becoming significantly more powerful.
What OpenAI Said in Response
OpenAI rejects the idea that model development proceeds independently of risks.
A company spokesperson told Reuters that OpenAI ensures model capabilities do not exceed the level the company can safely control and protect.
According to the company, when necessary, it pauses training or delays model releases.
OpenAI also said it is strengthening the security of research and testing infrastructure, expanding independent model evaluation, and developing real-time monitoring of model behavior.
Why Robinson's Departure Matters
This is not the first time safety experts have left leading AI companies and publicly warned about the risks of the race for more powerful models.
But in this case, the criticism comes from someone directly involved in preparing internal risk assessment rules and reports that accompanied OpenAI's biggest launches.
Robinson himself acknowledges that his conclusions remain his own assessment. He also does not claim that AI development must be stopped.
His main thesis is different: as model capabilities grow, the cost of error rises, so a system designed for constant experimentation and subsequent fixing of discovered problems may no longer be sufficient.
After leaving OpenAI, Robinson intends to work on AI safety issues from outside the company and push for stronger incentives for cautious development.
Based on: Reuters, The Atlantic