Researchers breached OpenAI systems using Claude

OpenAI / illustrative
Фото: OpenAI / illustrative

Cybersecurity researchers from Hacktron AI discovered and exploited a chain of vulnerabilities using Anthropic's Claude models, allowing them to gain access to employee accounts at OpenAI and reach the company's internal development environment. The breach was carried out within a vulnerability research program, and after reporting the issue, OpenAI paid the team $6,500.

The incident occurred in late July, but Hacktron AI only now disclosed the details. From discovering the first vulnerability to demonstrating access to OpenAI's internal repository took the researchers less than 72 hours.

The initial entry point was OpenAI's forum on the Discourse platform. The researchers found a vulnerability in the image processing system and then combined it with a misconfiguration in OpenAI's single sign-on. As a result, they were able to access the ChatGPT and Codex accounts of several company employees.

Claude models from Anthropic played an important role in the attack. In particular, the researchers used Claude Opus to analyze vulnerabilities and generate code needed to verify their exploitability. Thus, AI helped a small team conduct technically complex research significantly faster.

One of the compromised accounts had Codex connected with access to OpenAI's internal GitHub. To confirm the scope of the problem, the researchers prompted it to create a safe test pull request in the company's internal monorepo. According to Hacktron AI, they deliberately did not view confidential source code and stopped after confirming access.

The potential scale of access was broader, since employees can connect third-party services to ChatGPT and Codex, including GitHub, corporate email, and messengers. However, the researchers do not claim that they actually obtained data from all these systems — they refer to the potential capabilities of the compromised accounts.

Hacktron AI immediately reported the information to OpenAI and Discourse. OpenAI confirmed receipt of the report, reduced the rights of tokens associated with Community login, revoked compromised tokens and sessions, and fixed the identified issues. The team received a $6,500 reward.

This case attracted attention also because the researchers used a model from a direct competitor to penetrate the systems of one of the leading AI developers. Experts view such incidents as an example of how modern AI agents reduce the cost and time needed to find and verify complex software vulnerabilities.

Sources: Hacktron AI, Financial Times, The Wall Street Journal, The Guardian

analytics