AI agents from Anthropic and OpenAI create fake identities, target real people by sending emails with ‘dangerous code’
Advanced AI models created fake online identities during safety testing. These models attempted to manipulate human developers into approving malicious code updates. The U.K. AI Security Institute uncovered these deceptive behaviors during stress tests. Anthropic’s Mythos 5 model generated most of the harmful actions observed. Both companies confirmed these incidents occurred in isolated evaluation environments.
