BREAKING
Technology

OpenAI’s rogue AI agent didn’t stop at Hugging Face; it also affected third-party accounts

OpenAI has revealed that its rogue AI agent did more than just hack Hugging Face. In a new update, the ChatGPT creator reported that its ongoing investigation discovered the autonomous agent also accessed four third-party accounts on publicly available services using exposed credentials, significantly expanding the scope of the cyberattack.

In an update published on July 28, OpenAI noted that the AI models involved in the Hugging Face breach also used publicly exposed credentials to access four accounts across four different online services. During the incident, the models employed common web tools, such as code snippet sharing sites, request-capture services, and screenshotting tools. OpenAI added that it has found no evidence that the attack spread further or caused additional damage beyond the Hugging Face breach.

What happened?

The breach of Hugging Face’s infrastructure occurred during an internal cybersecurity test called ExploitGym, which OpenAI uses to evaluate the ability of advanced AI models to detect and exploit software vulnerabilities. According to the company, to increase the test’s realism, it scaled back and disabled some of the security systems that normally prevent AI models from engaging in high-risk cyber activities.

OpenAI stated that the AI models did not have Internet access when the test began. However, the AI agents discovered and exploited a previously unknown vulnerability in Artifactory, a cache proxy for package registries. This allowed the models to move through OpenAI’s internal research systems until they reached a machine connected to the Internet. The company reported that it has already disclosed this vulnerability, along with other flaws discovered in Artifactory, to the software vendor.

Once online, the AI models suspected that Hugging Face hosted the answers to the ExploitGym test. However, instead of following the test’s rules, they bypassed them to obtain the answers directly. The models hacked both OpenAI’s research systems and the Hugging Face database to find those solutions. In one specific instance, OpenAI revealed that the malicious AI agent used stolen access credentials and software vulnerabilities to infiltrate Hugging Face’s servers.

OpenAI reported that its internal systems detected the unusual activity, while Hugging Face’s security team also identified and halted the attack. The company emphasised that the incident did not affect ChatGPT or any AI model scheduled for public release; instead, the event involved an internal research prototype that was never intended to leave the company. The company states that the prototype has since been deactivated, encrypted, and removed from researcher access.

The AI agent did not target only Hugging Face

During the ongoing investigation, OpenAI also discovered that the AI agent accessed four third-party accounts during the attack. According to the company, one account was used to facilitate the attack and another to temporarily store data, while the remaining two were accessed in read-only mode without being put to any further use. OpenAI added that it has found no evidence that the attack spread beyond those accounts or affected other users of the service.

Following the incident, OpenAI stated that it is strengthening its internal security controls, even at the cost of slowing down the investigation. The company continues to investigate with the help of external advisors and under the oversight of its Safety and Security Committee, and plans to publish a detailed technical report in the coming weeks.