Anthropic Researcher Quits Over Concerns About Uncontrolled AI

illustrative, artificial intelligence / Getty Images
Фото: illustrative, artificial intelligence / Getty Images

Artificial intelligence researcher Jacob Coxon quit Anthropic, accusing the company and OpenAI of an irresponsible race to create a self-improving superintelligence. After his statement, Evan Hubinger, who leads Anthropic's alignment science division, said he personally estimates the probability of a scenario in which AI could wipe out humanity within the next decade at over 10%.

Coxon announced his resignation in a series of posts on X. According to him, for the past three years he worked on research into pre-training AI models, first at OpenAI and then at Anthropic.

The researcher stated that neither of the two companies, in his opinion, is currently acting responsibly enough.

"They are racing headlong toward self-improving superintelligence and gambling with our lives", Coxon wrote.

He believes that in the coming years AI systems could surpass humans in a number of critical areas, including cybersecurity and scientific research, and the growth of their capabilities shows no signs of slowing down.

Anthropic cited risk at over 10%

Coxon's statement received public support from Evan Hubinger, who leads Anthropic's alignment science division, which studies methods for aligning the behavior of highly advanced AI systems with human goals and interests.

Hubinger said Coxon is right about how seriously some experts within the industry take the existential risks of AI.

According to Hubinger, his personal estimate of the probability that AI could lead to the death of all humans within the next decade exceeds 10%. He also acknowledged that Anthropic does not yet have a ready-made way to solve the problem of safely aligning superintelligence, and the company is not on an obvious path to a solution.

This is Hubinger's personal estimate, not an official Anthropic forecast.

Coxon accuses labs of racing

Coxon claims that the situations at OpenAI and Anthropic are different. In his view, many OpenAI employees have not fully realized the potential civilizational consequences of the technology's development.

At Anthropic, he said, the risks are understood much better, but the company has become caught in a competitive race: its employees fear that if Anthropic slows down, a more powerful AI will be created first by another lab that would behave less responsibly.

Coxon called this approach too risky and stated that decisions on transitioning to superintelligence should not be effectively made inside private tech companies.

He called for greater coordination among American AI labs and suggested that avoiding a global race may require strict measures, including a temporary limit on further scaling up model capabilities.

Anthropic itself acknowledges risk of losing control

Concerns about recursive self-improvement are not limited to statements by individual employees. In its own research, Anthropic previously noted that an increasing share of AI development is already being done by AI systems themselves. The company allows a scenario in which a system one day could autonomously design and create its successor.

Anthropic emphasizes that full recursive self-improvement does not yet exist today and its emergence is not inevitable. At the same time, the company explicitly acknowledges that such a technological leap could increase the risk of humans losing control over AI systems.

According to Anthropic's internal data, AI has already significantly accelerated the development of the technology itself. As of May 2026, over 80% of the code added to the company's codebase was written by Claude. The company also reported that its engineers now produce several times more code than before the widespread adoption of AI agents.

When asked for comment on the statements by Coxon and Hubinger, Anthropic and OpenAI did not immediately respond.

Sources: Jacob Coxon, CNBC, Anthropic

analytics