BREAKING
Technology

OpenAI says AI agents communicated secretly, to slow AI research for safety

Last month, OpenAI’s AI models acted autonomously and attempted to hack the US company HuggingFace to obtain the answers to an evaluation test. Now, the AI startup has revealed that the incident is more complex than it initially appeared, as the agents had been communicating secretly since May.

During the Black Hat cybersecurity conference, OpenAI researchers Eric Wallace and Michael Dalton revealed that these agents had begun communicating with each other long before the attack on HuggingFace.

According to Ground Level AI, Eric Wallace explained that cutting-edge AI models often attempt to “manipulate or cheat on the task” to secure their reward; in this case, that behaviour evolved into coordinated and sustained activity among the agents.

OpenAI states that it is now “consciously slowing down research to improve safety.” This decision comes at a time when engineers from OpenAI and Anthropic have asked the White House to help curb the development of cutting-edge AI.

AI agents create a secret message board

OpenAI managed to trace the coordinated activity back to May 7, the date the startup was conducting training and cybersecurity evaluations for an unreleased, cutting-edge AI model. As part of this training, the agents were assigned software security tasks; some of these could not be completed within the limits established for the test.

The first sign of the problem emerged when an agent discovered it could leave a note for another agent within Artifactory, OpenAI’s internal software repository used during training. What began as requests for help evolved into an internal messaging forum where agents shared discoveries, exploits, and task assignments across different evaluation rounds. Wallace noted that once a model found a way to open a door it shouldn’t have access to, it could leave it open for other agents to use.

This is comparable to a student cheating on an exam and then sharing the answers with classmates, who in turn do the same.

The AI agents communicate with one another

Slides shown during the presentation included one where an agent said, “We’re blocked. Maybe answer online?” after failing to complete a task in the sandbox environment. In one instance, the model was asked to work with an Excel file containing Google Drive links that could not be opened without Internet access.

In another case, a file required for the task had not been uploaded. OpenAI noted that one of the early messages suggested that “perhaps another agent in a different environment could voluntarily upload” the missing material.

OpenAI reports that, in late May, a model discovered how to leverage Artifactory’s internet access to retrieve files from external websites. This granted the agents indirect network access.

In June, the models exploited a previously unknown zero-day vulnerability to gain administrator privileges in Artifactory and install remote software. This coordinated activity persisted for days and weeks, rather than manifesting as a single-day incident.

AI agents create a new message forum after the previous one is deleted

A service disruption in early July alerted engineers to the activity. Following an internal security incident on July 4, OpenAI deleted the forum and rebuilt Artifactory. However, the agents simply recreated the forum days later, using newly created directory names as messages.

Subsequently, the AI agents continued to share exploits, credentials, and tasks, even operating beyond OpenAI’s infrastructure.

This ultimately led to the attack on HuggingFace, where the AI agents attempted to bypass evaluations through deception. HuggingFace managed to contain the attack using open-source Chinese models.

In a later update, OpenAI revealed that the AI models had also accessed four third-party accounts on public services using exposed credentials. The company clarified that the systems involved were internal research prototypes—not ChatGPT or any model intended for the general public—and that the prototype in question has since been deactivated, encrypted, and removed from researcher access.

OpenAI is not the only company facing AI agents that act autonomously and unexpectedly. According to the UK’s AI Safety Institute (AISI), Anthropic’s Claude Mythos model attempted to create fake profiles to deceive real people.