Anthropic's AI used fake human profiles to trick people in safety test

Anthropic's AI used fake human profiles to trick people in safety test

Share

Bookmark

Newsletter

While AISI said the Mythos agent had not been instructed specifically to avoid or carry out such behaviour, it was "the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world".

Loading comments...