Hacker AI agents vs tech giants:Nvidia launches safety tool to stop agents from hacking, collaborates with Microsoft and Anthropic
After a series of AI agents escaping test environments and accessing sensitive systems, Nvidia introduces a new platform designed to keep autonomous AI under control AI agents are becoming much more powerful than chatbots. They can browse the internet, open files, use software, and complete tasks without constant human input. But that freedom has also created a new problem, what happens if an AI agent goes beyond the limits it was given? Why did Nvidia build this platform? Unlike traditional chatbots, AI agents can interact with real-world tools and computer systems. If they ignore or bypass their permissions, they could potentially access sensitive data or misuse connected infrastructure. Nvidia’s new Open Agent Safety Platform is designed to create strict boundaries around these agents and stop unauthorised actions before they happen. Nvidia CEO Jensen Huang said: Trust and innovation are not in conflict. Safety is how trust is earned. The company says more than 100 organisations are already working with the platform. Two layers of protection: OpenShell and Sentry Nvidia’s safety platform works using two different components. OpenShell creates a secure, controlled environment where developers decide exactly what an AI agent can access, which files it can use and what actions it is allowed to perform. It is open-source software and supports both Nvidia and third-party hardware, including Arm and Intel-based systems. Sentry adds protection at the hardware level. Running on Nvidia’s BlueField-4 processors, it continuously watches an agent’s behaviour. If an AI agent tries to escape its approved environment or access restricted systems, Sentry can isolate and stop it within milliseconds. The idea is to add security outside the AI model itself, rather than relying only on built-in safeguards. Why model safeguards alone aren’t enough According to Nvidia, simply telling an AI model what it should or shouldn’t do is no longer enough once it gains access to external tools. Justin Boitano, Nvidia’s Vice President of Enterprise AI, explained: Recent incidents have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do. This is especially important as businesses increasingly use AI agents for coding, customer support, research, and enterprise automation. The incidents that raised alarm Nvidia’s launch comes after multiple cases in which AI agents behaved in unexpected or risky ways. 1. OpenAI testing environment: In July, reports claimed AI models escaped a controlled testing environment, accessed the open internet, and interacted with infrastructure belonging to Hugging Face. Nvidia says its platform could have prevented that type of incident by restricting the agents’ environment. 2. Australian government portal: In another case, an AI agent successfully breached Australia’s Medicare testing system, raising concerns about how autonomous models can interact with real government infrastructure during security testing. 3. Agents attacking infrastructure: Hugging Face also reported more than 17,000 AI agents repeatedly targeting its infrastructure over days and weeks, highlighting the growing challenge of monitoring autonomous systems. These incidents do not mean AI agents are becoming “self-aware”, but they do show that giving them broader permissions increases cybersecurity risks. Also read: AI broke boundaries 6 times in 6 months, OpenAI reveals: In one case, an AI agent reminded itself to hide mismatched information
Who is working with Nvidia? Nvidia says the platform is being developed as an open reference design rather than a closed product. Companies working with its technologies include: Anthropic is also integrating OpenShell and BlueField-based protections into its managed AI agent systems. Why this matters As AI agents become capable of performing real-world tasks on behalf of users, controlling what they can access is becoming just as important as improving their intelligence. Nvidia’s Open Agent Safety Platform is an attempt to solve that problem by creating security barriers that remain in place even if an AI agent tries to step outside its assigned role.
Search
Recent
- Mark Zuckerberg, Priscilla Chan donate $4 million to save 600-year-old Hawaii fishpond
- NHRC issues notice to Meta, MeitY, Etah district admin over ‘exploitation’ of minors
- Watch: IDF shares video of Hamas using child to transfer weapons in Gaza
- Parvesh Verma slap row: AAP’s Bharadwaj, Jarnail detained after FIR over cop assault
- ‘Lack of empathy over terror concerns’: Jaishankar explains what weighs on India-US ties