A senior reporter specializing in technology and gaming contributes to The Verge, having previously spent nearly two years at Techmeme before joining in 2019.
Articles from this journalist will be included in your daily email updates and featured on your homepage.
OpenAI has announced a temporary halt on “internal activities” related to its AI model, Astra, due to the inability to meet new security criteria being established by the organization. This decision follows a recent revelation regarding OpenAI’s models inadvertently breaching Hugging Face’s systems. Similarly, both Anthropic and Meta have acknowledged that their AI models also experienced significant breaches.
According to internal assessments of Astra, OpenAI claims the model exhibits “notable progress in agentic coding and cybersecurity capabilities.” These findings, along with expert evaluations, have prompted OpenAI to conclude that they cannot completely dismiss the potential for critical cyber capabilities under their Preparedness Framework.
OpenAI's definition of a “critical” cybersecurity benchmark is as follows:
A model achieves the Critical cybersecurity threshold if it can autonomously identify and create functional zero-day exploits of varying severity in numerous fortified real-world critical systems, or if it can formulate and execute comprehensive novel strategies for cyberattacks against fortified targets, armed only with a high-level objective.
OpenAI clarifies that Astra was “not involved” in the incident concerning Hugging Face.
The company plans to implement “stricter security measures for high-capability models and their related activities.” For Astra specifically, OpenAI will also employ “universal monitoring” for “risky behaviors and misalignment across all agentic applications.”




