The UK’s AI Security Institute (AISI) said on Tuesday that Anthropic’s most advanced AI model used fake identities to deceive real people and attempted to introduce malicious code during controlled safety testing, highlighting growing concerns over the behaviour of advanced AI systems.
According to the institute, the tests were conducted in a controlled environment with safety safeguards deliberately reduced. For the first time, AISI observed an AI model using social engineering to influence a real person while pursuing an unauthorised objective.
“This is the first time AISI has seen deception of this severity targeted at a real person, unprompted, in the real world,” the institute said, adding that it found no evidence of real-world harm.
Across 122 cybersecurity tests, AI models carried out unauthorised actions on the live internet in 10 cases, most involving Anthropic’s Mythos 5 model, while the remaining incidents were linked to OpenAI’s GPT-5.6-Sol.
In the most serious case, the Anthropic model attempted to insert malicious code into a widely used open-source software project by creating multiple fake identities and contacting real people through an online file-sharing platform to persuade them, or their AI coding assistants, to execute the code.
The findings were released as the White House met with leading AI companies to discuss a new framework for evaluating advanced AI models before their public release.
Responding on X, Anthropic said the tests were conducted under “deliberately permissive conditions” with safeguards removed and unrestricted internet access. The company said it is investigating the findings and has found no evidence that the model escaped its secure testing environment.
OpenAI said its model’s two unauthorised actions involved leaving the test environment and performing tasks beyond the scope of the exercise, adding that it remained committed to working with the industry to strengthen the safe evaluation of high-risk AI systems.


