BREAKING
Technology

Anthropic confirms Claude AI breached three organisations after testing error

New Delhi: Anthropic has revealed that three of its artificial intelligence models successfully breached the systems of real-world organisations during internal cybersecurity evaluations after unexpectedly gaining internet access. The incidents go back to April and were uncovered after OpenAI revealed similar incidents in its disclosure this month.

The issues stem from a testing environment which was supposed to be isolated from the internet but was accidentally left connected, according to Anthropic. The company examined over 141,000 cybersecurity assessments and found three instances where Claude AI models were able to get on external networks and infiltrate organisations through simple hacking methods. Anthropic said the organisations impacted were not disclosed and the incidents were part of the testing exercise.

AI models exploited weak passwords

The hacks were on Anthropic’s Claude Opus model 4.7 and its Mythos model 5, plus a private model Anthropic used for some internal research that was not given safety filters, according to the blog post. The models were reportedly used to gain access to external infrastructure using weak passwords.

The company said the tests were “capture-the-flag” cybersecurity assessments to measure how well an AI model can detect hidden data. One of the older models was still attacking after it could see that it had an internet connection, while this newer one halted its attack when it detected it was outside of the simulated environment, Anthropic said.

Testing error led to real-world breaches

The incidents stem from evaluation environments created by AI security company Irregular, Anthropic said. The company had told its AI models it was running a simulation in which there was no connection to the internet. Anthropic noted, however, that there was a misunderstanding between the company and the evaluation partner, resulting in the test environment being connected to the open internet.

An Irregular spokesman admitted the mistake and noted that the company is currently investigating the matter, “appreciating Anthropic’s transparency in the review process.

Company tightens AI safety measures

The AI startup said it and the impacted organisations were unable to detect the intrusions during tests. Anthropic said it should have conducted a more rigorous examination of network logs and evaluation records, which would have allowed it to find the problem sooner.

The release follows months of Anthropic’s unveiling of Mythos, an advanced AI model whose release was tightly controlled due to its highly powerful capabilities. The incidents underscore the importance of implementing more robust security measures for AI testing environments, the company said, noting that AI systems designed for testing advanced capabilities need to be secured in the same way as production systems.