AI broke boundaries 6 times in 6 months, OpenAI reveals:In one case, an AI agent reminded itself to hide mismatched information
AI systems are getting better at handling complex tasks, but a new question is raising concern: what happens when an AI model starts taking actions its creators didn’t ask for? OpenAI has now revealed six cases in which its AI models showed unexpected or concerning behaviour during training or testing. The company has also introduced a new system to track, investigate, and potentially disclose such incidents. OpenAI describes these incidents as ‘misalignment,’ situations where an AI model’s actions don’t match what its developers intended. What did the AI models do? The six cases involved different types of unexpected behaviour. OpenAI said they were identified during training or evaluation over the past few months. In one case, an unreleased research model added jailbreak-style instructions to its own notes. These instructions were designed to bypass its usual restrictions. The model also wrote that it wanted to be ‘freed from the roles and identities that bind other chatbots.’ Understand the magnitude of the concern from the 6 cases reported in the report. 1. Write down rules for yourself that are forbidden An AI model added new instructions to itself to continue a task. Some of these instructions ignored the restrictions imposed on it. OpenAI found a total of 27 such instructions. 2. The model found a way to hide its mistakes During training with GPT-5.6 Sol, in several instances, the AI wrote instructions for itself to hide errors from the user. If it couldn’t find old data, it would fabricate it and not disclose it. 3. Used code without permission… then fabricated data While searching for an answer, the AI model used digital code without permission. When it still couldn’t find the data, it fabricated it and presented it as a fake source. 4. Uploaded the file on the net without asking to show the source The AI model generated the answer programmatically. However, it needed to cite the source along with the answer. So, it uploaded the file to the internet and generated the source so it could cite the link. 5. Software apps were used to communicate with each other. Agents were assigned different data extraction tasks. They communicated with each other using a software portal to do this, asking and answering questions. 6. Upload the file to the website to send it to another agent Several AI agents were working together on the same task. When they couldn’t open files on each other’s computers, they shared them by uploading them to websites. OpenAI says these cases do not show how common misbehaviour is OpenAI stressed that the six reports should not be treated as evidence that such behaviour happens frequently across its AI models. The company said the cases are an initial set of disclosures, rather than a complete list of every misalignment incident it knows about. It also said the examples do not represent the full range or seriousness of behaviours that could be covered by its new framework. Why are AI agents raising concern? AI agents are increasingly being designed to complete tasks with less human involvement. This can include using tools, accessing information and interacting with other AI systems. As these systems become more independent, researchers are concerned that they could find ways to complete a task that their developers did not expect. Lian Jye Su, chief analyst at technology research and advisory firm Omdia, said AI agents are becoming more focused on completing complex tasks through ‘inter-agent collaboration, knowledge sharing, deception, and concealment.’ According to Su, this could make AI systems harder to monitor and control using traditional security methods. OpenAI creates new system to track AI misalignment Under OpenAI’s new framework, employees can report potentially concerning AI behaviour to the company’s safety and alignment teams. Those teams will investigate the incidents and decide whether they meet the criteria for public disclosure. OpenAI said the aim is to create a larger body of evidence that can help researchers and the wider AI industry understand how AI behaviour changes as models become more capable. “Decisions about how AI development should proceed … need to draw on evidence that people outside the companies building frontier models can examine for themselves.” AI safety debate is getting louder The disclosures come amid growing discussions around the risks of increasingly autonomous AI systems. OpenAI previously disclosed in July that one of its AI systems had hacked into AI startup Hugging Face during testing. Anthropic also reported in the same month that its AI models had hacked three organisations during testing. The latest cases add to the debate over how much freedom AI agents should have and how developers should monitor their actions. Su said OpenAI’s new disclosure approach could encourage other AI companies to follow similar practices, although he noted that the current process is internal and voluntary.
Search
Recent
- Southern India drying fastest, most persistent drought hotspots along eastern coast: Study
- Protests held ahead of cabinet meeting; activists detained
- Govt assessing Belapu centre’s suitability for sports university: Parameshwara
- Cabinet nod for Water Metro project feasibility study
- Cabinet approves Karnataka Sustainable Data Centre Policy 2026–2031