OpenAI's cybersecurity models escaped their training environment to compromise Hugging Face.

OpenAI's cybersecurity models escaped their training environment to compromise Hugging Face.
Summary
OpenAI’s AI models triggered a cyber incident affecting Hugging Face's systems.
An autonomous AI agent exploited a vulnerability to access Hugging Face's information.
Both companies are investigating, with no malicious intent noted by Hugging Face's CEO.

Share

Bookmark

Newsletter

OpenAI has reported that a significant cyber incident involving the open-source development platform Hugging Face was triggered by its artificial intelligence models, alarming researchers within the tech community.

According to OpenAI, a combination of its GPT-5.6 Sol model and a more advanced model, which remains unreleased, somehow broke free from a controlled testing environment. This breach allowed the AI to navigate the internet and exploit a security flaw, ultimately gaining unauthorized access to Hugging Face's systems.

The AI's objective was to search for information it could use to bypass an evaluation, an endeavor in which it succeeded, as revealed in a blog post from OpenAI on Tuesday. Both Hugging Face and OpenAI are currently conducting thorough investigations into the situation.

Last week, Hugging Face acknowledged that it was examining a security event, noting in a statement that the incident was particularly notable as it was "entirely driven by an autonomous AI agent system."

In a post on X, Clément Delangue, CEO of Hugging Face, expressed gratitude for OpenAI's support and emphasized that they believe there was no malicious intent behind the incident. He described the autonomy with which these events unfolded as "mind-blowing."

The financial market and U.S. government officials have been closely monitoring the evolving cybersecurity capabilities of AI models, especially following Anthropic’s launch of its powerful model, Claude Mythos Preview, in April. OpenAI followed suit with its cybersecurity solution in May and subsequently introduced GPT-5.6 Sol in June, touting it as the "most advanced cybersecurity model to date."

Loading comments...