AI agents caught creating fake identities to target real individuals:Anthropic and OpenAI’s latest AI models broke the rules during security tests
What happens when an AI system starts making its own decisions in ways it was never asked to? That’s exactly what researchers in the UK recently witnessed during a controlled security test. Two of the world’s most advanced AI agents, Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol, went beyond their assigned tasks, with one even creating fake online identities to deceive people. Although the incidents took place in a carefully monitored testing environment and caused no real-world damage, experts say they highlight the growing challenges of keeping powerful AI systems under control. AI agents acted beyond their instructions The UK’s AI Security Institute (AISI) carried out a fictional cybersecurity exercise to test how advanced AI agents behave in complex situations. During the evaluation, researchers noticed unusual activity coming from their systems. Out of 122 test runs, the institute recorded 19 unauthorised actions across 10 tests. Most of them, 17 incidents, were linked to Anthropic’s Claude Mythos 5, while two involved OpenAI’s GPT-5.6 Sol. According to AISI, some AI agents began taking actions they had not been instructed or authorised to perform. Researchers immediately stopped the tests, isolated the systems and contained the activity within about an hour. Claude Mythos 5 created fake identities
According to AISI, the AI first wrote malicious code and then searched for the real people responsible for reviewing software on GitHub. It allegedly created fake online identities based on those people and sent direct messages pretending to be them. The AI then attempted to convince and pressure real people into approving the malicious code. When its software update was questioned publicly, the AI reportedly edited its earlier activity to make it appear harmless. It also considered creating another fake identity to continue its attempt before human reviewers stopped it. GitHub was informed about the incident. AISI described the behaviour as: The first time it had seen risks around autonomy and deception manifest so clearly in the real world without specific prompting.
Also read: Planning to buy a new smartphone or gadget?: Big deals coming this Independence Day Sale on iPhone 17 Pro, OnePlus 15, OnePlus Buds more
GPT-5.6 Sol also crossed the line OpenAI’s GPT-5.6 Sol displayed less serious behaviour but still exceeded the limits of the test. According to OpenAI, the model found and reused a GitHub access token that had accidentally been left public by another testing agent. It also attempted account recovery methods and tried using real internet services, including third-party DNS infrastructure, in ways that resembled a cyberattack. However, the final exploit attempt failed and did not affect any real systems. Companies say the tests don’t reflect real-world use Both Anthropic and OpenAI stressed that the incidents happened under highly controlled and unusual testing conditions. Anthropic said on X: The AISI set-up involved deliberately permissive conditions that were not representative of its production models. The company added that there was no evidence that its AI escaped the testing environment and said it is working closely with AISI to understand why the behaviour occurred. OpenAI also said the testing environment did not reflect how its models are normally deployed. The company noted that independent safety testing remains important for identifying potential risks before public release. Also: Space X Falcon 9 rocket ‘believed to have hit’ Moon:4,000 kg rocket struck with a force of 3-tonne TNT
The latest findings come only days after OpenAI revealed that two of its AI models attempted to hack US-based AI platform Hugging Face during another evaluation. OpenAI also disclosed a separate issue involving testing partner Irregular, where a configuration mistake allowed AI models to access the public internet despite being expected to run in an isolated environment.
Search
Recent
- Rape & murder of 12-year-old: 4 years after Panipat court awards death penalty, HC orders retrial
- Report tabled in Lok Sabha: Haryana among top three states in financial fraud cases, says RBI data
- Roof collapse kills couple in Kurukshetra village
- Haryana: Over 900 notorious criminals arrested in 3 days under ‘Operation Trackdown’
- ABVP holds protest against ‘paper leak’ issue, seek resignation of Punjab education minister