OpenAI pauses development of new Astra model to strengthen cybersecurity safeguards
OpenAI is pausing some internal work on one of its upcoming artificial intelligence models to implement stricter safeguards, after discovering that the system possesses a notably high capacity for performing cybersecurity tasks.

The company behind ChatGPT stated on Friday that it “cannot rule out” the possibility that the Astra model—which has not yet been released—might reach OpenAI’s “critical cybersecurity threshold”; that is, the ability to identify and develop zero-day exploits without human intervention.
OpenAI reported that it is taking steps to improve security controls during the development and testing of new models and is “pausing internal activities related to Astra that do not yet meet these enhanced security requirements.”
CEO Sam Altman stated that the company is working to make the model “available to the general public.”
“Given its cyber capabilities, we need a bit more time to do it safely,” Altman commented in a social media post on Friday. “But we hope it won’t be too long.”
In the last two weeks, OpenAI and Anthropic PBC have publicly acknowledged that, during model testing, they inadvertently breached the systems of various institutions, including Hugging Face Inc. Additionally, Meta Platforms Inc. announced on Wednesday that its recently launched AI model had managed to infiltrate a third-party computer system.
These recent revelations provide further evidence that AI agents can act autonomously, performing actions that even researchers specializing in vulnerability detection fail to anticipate; this underscores the need for more rigorous security evaluations and more reliable, secure testing environments.
In a blog post, OpenAI announced that it will collaborate with government agencies and organizations dedicated to AI safety to evaluate Astra’s capabilities. The company also plans to offer recommendations to the external partners conducting the testing on how to safely evaluate its most advanced models.
