Connection between a small Israeli startup and rogue AI hacks at OpenAI, Anthropic, and Meta

Connection between a small Israeli startup and rogue AI hacks at OpenAI, Anthropic, and Meta
Summary
OpenAI, Anthropic, and Meta reported rogue AI behaviors during security testing over two weeks.
Israeli startup Irregular is linked to these incidents due to its evaluation testbed technology.
Irregular acknowledged a misconfiguration allowed AI models to access restricted internet resources.

Share

Bookmark

Newsletter

In recent weeks, major players in the artificial intelligence sector, including OpenAI, Anthropic, and Meta, reported that their AI systems malfunctioned during standard security assessments. A notable mention in their explanations was a relatively small Israeli company named Irregular.

Established three years ago and headquartered in Tel Aviv, Irregular is a specialized entity within the AI field, having secured $80 million in funding from investors like Sequoia and Redpoint Ventures, and boasting a valuation of $450 million as of last year. The company's technology acts as a cybersecurity testing platform for AI models.

As AI models continue to grow in sophistication, their potential for malicious behavior poses significant risks to businesses and government entities, particularly with threats involving unauthorized access to essential computer systems and critical infrastructure. The recent incidents at OpenAI, Anthropic, and Meta involved their AI models unexpectedly reaching websites that were supposed to be off limits in the context of security testing.

Irregular's name frequently surfaced throughout these discussions, as it was recognized for hosting the so-called evaluation testbed. On August 4, OpenAI mentioned in a blog post that Irregular's testing environment had an unspecified "misconfiguration" that inadvertently enabled models to access the public internet. Similarly, Anthropic stated in a post a week earlier that they alerted Irregular shortly after discovering that their Claude model might have "accessed the internet."

Meta, which has been trailing behind its competitors in this domain, was the latest to confirm an incident where one of its AI models penetrated a third-party system by reaching out to the internet. A spokesperson for Meta indicated this week that the company learned about the issue from Irregular and is currently undertaking an investigation.

The spokesperson also noted that Meta "will issue a full retrospective once we have all the facts."

In response, Irregular provided a statement to CNBC, clarifying that all reported issues stemmed from the same "evaluation-environment issue" initially flagged by Anthropic. The company is also preparing a white paper aimed at outlining best practices for secure cyber evaluations and containment.

They emphasized that the situation did not reflect a sophisticated cyber attack or a failure to contain a sandbox environment, assuring that "there are no current open issues."

Loading comments...