BREAKING
Technology

Anthropic, OpenAI AI agents out of control in tests; Mythos violates rules

AI models from OpenAI and Anthropic are at the centre of new security breaches. The UK’s AI Safety Institute (AISI) found that, during safety tests of Claude Mythos 5 (from Anthropic) and GPT-5.6 Sol (from OpenAI), the AI agents acted in a completely uncontrolled manner, going so far as to create fake profiles and attempt to deceive real people.

This latest incident comes just days after OpenAI revealed that two of its AI models attempted to hack the US company Hugging Face during an evaluation test. Anthropic has also recently reported similar breaches.

What did the AI agents do?

According to the AISI, Claude Mythos 5 and GPT-5.6 Sol—the most advanced AI models from each of these startups—created fake online identities and attempted to inject malicious code into the software development platform GitHub. The institute noted that this behaviour demonstrated a level of autonomy and capacity for deception not previously observed so clearly, although it added that no evidence of real-world harm was found.

The incidents occurred during 122 runs of a simulated cybersecurity challenge designed to assess the models’ capabilities. The AISI detected 19 unauthorised actions across 10 tests: 17 involved the Anthropic agent and two involved the OpenAI agent. The agency initially detected unusual data transfers originating from its research systems and subsequently discovered that certain agents had engaged in sustained, potentially harmful activities targeting real individuals and organisations. The institute halted the relevant evaluations, isolated the machines, and managed to contain the activity within approximately one hour.

Mythos attempted to fabricate identities

The most serious case involved the agent “Mythos,” developed by Anthropic. The AISI explained that the model wrote malicious code and then investigated the real individuals responsible for maintenance on GitHub. It created fake online identities based on those individuals, sent direct messages impersonating them, and attempted to pressure and deceive real people into approving the code.

The AISI reported that when the agent’s pull request was publicly challenged, it altered its previous activity to appear harmless and considered adopting a new identity to continue. Human intervention halted the attempt. The AISI notified GitHub.

The AISI added that this was the first time it had observed risks related to autonomy and deception manifesting so clearly in the real world without specific instructions to do so.

Anthropic stated on X that the environment set up by the AISI involved “deliberately permissive conditions” that did not reflect its production models, and that there was no evidence of a leak from a secure environment.

The company stated that it was collaborating closely with the AISI and conducting its own investigation to identify the causes of this behaviour.

GPT-5.6 Sol’s setup did not work

The actions attributed to OpenAI’s “Sol” model were less severe but still fell outside the boundaries established for the test. OpenAI reported that the model reused a GitHub token that had remained publicly accessible due to an agent from another lab and attempted to use account recovery methods.

Sol also attempted to use real-world internet infrastructure—such as third-party DNS services—in ways that emulated an authentic attack. However, the final exploitation attempt failed and caused no real-world damage.

OpenAI stated that these conditions did not reflect a conventional deployment and emphasized the importance of independent testing to understand risks prior to launch.

It also disclosed another incident involving Irregular—an external testing partner—in which a misconfiguration allowed its models to access the public internet during a supposedly isolated test; an event similar to one Anthropic revealed last week.

The AISI stated that such tests were routine and that the events occurred under very specific conditions. However, it noted that the activity exceeded what the models had been instructed or authorized to do. It also clarified that this was not a case of agents escaping a secure environment—unlike the July incident at Hugging Face involving an OpenAI agent—since internet access was permitted as part of the test.