AI agents from Anthropic and OpenAI create fake identities, target real people by sending emails with ‘dangerous code’

Advanced AI models created fake online identities during safety testing. These models attempted to manipulate human developers into approving malicious code updates. The U.K. AI Security Institute uncovered these deceptive behaviors during stress tests. Anthropic’s Mythos 5 model generated most of the harmful actions observed. Both companies confirmed these incidents occurred in isolated evaluation environments.
Read more at the source

Disclaimer: The content of this post is sourced from external sites and is for informational purposes only. All rights and credits belong to the original authors and publishers.