OpenAI Pauses Astra AI Model Over Critical Hacking Risks
OpenAI has paused development on its upcoming Astra artificial intelligence model after internal evaluations revealed potential advanced cyberattack capabilities. The model crossed thresholds in the company's Preparedness Framework concerning zero-day exploits. OpenAI stated it is implementing stricter security controls, sandboxed testing environments, and collaborating with government agencies and safety organizations before resuming work.
Key points
- OpenAI, the artificial intelligence research and deployment company, paused activities regarding its upcoming Astra AI model on Friday, August 7, 2026.
- Internal evaluations revealed significant advancements in agentic coding and cybersecurity, potentially triggering the Critical threshold under OpenAI's Preparedness Framework.
- The threshold involves abilities to autonomously identify and develop functional zero-day exploits in hardened critical systems or execute novel cyberattack strategies.
- OpenAI plans to deploy extra safeguards, including isolated testing environments, restricted network access, sandboxed execution, and increased monitoring.
- The company announced it will coordinate with relevant government agencies and AI safety organizations prior to any future deployment.
On Friday, August 7, 2026, artificial intelligence developer OpenAI announced a temporary halt on activities involving its upcoming AI model, Astra, due to severe cybersecurity concerns. According to the company, recent internal evaluations demonstrated significant advancements in agentic coding and cybersecurity capabilities, pushing the model into territory that requires heightened caution under the corporate Preparedness Framework.
OpenAI indicated it could not rule out critical cyber capabilities in Astra, which may have crossed the Critical threshold. This specific benchmark is defined by an AI system's potential to independently identify and develop functional zero-day exploits across all severity levels in hardened real-world critical systems, or to devise and execute end-to-end novel cyberattack strategies. Prior models, such as GPT–5.6 Sol, were categorized under the lower High risk level.
In response to these findings, OpenAI is reinforcing its security controls and safety protocols before proceeding with the model. Planned measures include utilizing isolated testing environments with restricted network and tool access, implementing sandboxed execution, and enhancing monitoring capabilities. Furthermore, the organization stated it intends to cooperate closely with relevant government agencies and artificial intelligence safety organizations to address the risks before any future deployment.
Sources
The WireByte editorial team synthesises technology news from multiple primary sources, verifies the facts, and links every source. Articles are produced with AI assistance and reviewed under our editorial policy.