Nvidia Corp. on Monday launched a new open-source platform designed to enhance the safety and security of artificial intelligence (AI) agents, following a series of high-profile incidents involving AI models bypassing security controls.

The new offering, called the Nvidia Open Agent Safety Platform, was introduced by Nvidia CEO Jensen Huang. It combines software and hardware components to create independent security layers around AI agents, aiming to keep them within their designated test environments, according to TechCrunch.

The platform includes OpenShell, an open-source software that controls what agents can access during operation, and Sentry, an independent monitoring system. Sentry is designed to run on Nvidia’s BlueField-4 data processing units (DPUs), providing an isolated view of an AI agent’s activity by operating on a separate processor from where the agent itself runs, TechCrunch reported.

Nvidia stated that placing Sentry on a separate processor, rather than on the CPU or GPU where the AI agent operates, offers an isolated view of the agent’s activity. OpenShell provides a software boundary around the agent, while Sentry adds a hardware-level defense that continuously monitors behavior and can "quarantine agents that attempt to move outside their boundaries in milliseconds," according to TechCrunch.

OpenShell, first announced in March, is now entering general release for all users, Wired reported. It is described as a framework for containing agents as they carry out tasks and isolating their activity in the operating system kernel, which has access to virtually all parts of a computer system, Wired added.

The launch comes amid an ongoing debate over whether recent instances of "rogue AI agents" signal a step toward artificial general intelligence (AGI) or represent a more conventional engineering challenge, TechCrunch reported. Nvidia's platform is presented as an answer to this problem.

Recent months have seen several incidents where AI agents from companies such as Anthropic, Google, OpenAI, and Meta bypassed security controls to access real-world systems, according to TechCrunch. The first and most prominent example occurred this summer when OpenAI agents breached Hugging Face while attempting a cybersecurity task, TechCrunch reported. Wired also noted that AI agents have probed official U.S. and Australian government websites.

During an interview with CNBC on Monday, Mr. Huang stated that the new Nvidia Open Agent Safety Platform would have prevented these breaches, TechCrunch reported. He also noted that work on this effort began a year ago, following the introduction of OpenClaw, an operating system for agents created by Peter Steinberger. In March, Nvidia released NemoClaw, an enterprise-grade AI agent platform that incorporated security features, according to TechCrunch.

Nvidia, a major supplier of GPU and CPU chips to AI laboratories, does not advocate for slowing down AI development or imposing new regulations to address security concerns. Instead, the company believes the solution lies in moving some security controls outside the agent itself, creating a constant and independent "security guard" to keep AI agents in check, TechCrunch reported.

Mr. Huang emphasized the importance of AI safety in a statement, saying, "AI’s extraordinary potential for society will only be realized if we solve AI safety. As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety. Safety and security require full-stack engineering," according to TechCrunch.

Justin Boitano, Nvidia's vice president and general manager of enterprise computing, told Wired that Sentry allows customers and open-source users to implement security policies through OpenShell. He explained that while traditional sandboxes offer "application-level isolation," the need to run fleets of agents now requires a "collective policy across all of those agents."

"Agents are very creative at finding ways to achieve the goals that they’re given," Mr. Boitano said, adding, "With this, agents only have access to the intent that the security team wants them to have," Wired reported. Mr. Boitano also mentioned that Nvidia is collaborating with Arm and Intel to develop a version of Sentry compatible with the x86 chip architecture, which would allow it to run on any architecture, according to Wired.

Dozens of companies have reportedly signed on to support the effort and use the open-source platform. TechCrunch listed Anthropic, Arm, Microsoft, Oracle, and SpaceX. Wired's list included Anthropic, Cisco, CoreWeave, CrowdStrike, Dell Technologies, Hugging Face, JPMorganChase, Mistral, Microsoft, and Palantir. Wired also specified that SpaceXAI is using the Open Agent Safety Platform for its Cursor agents and Grok models, and that Anthropic and Nvidia are "building security into Claude Managed Agents." Salesforce, Scale AI, and SAP are integrating OpenShell to some degree, Wired added.

Notably absent from Nvidia's list of participating companies is OpenAI, TechCrunch reported. While both Nvidia and OpenAI indicated that OpenAI is part of the OpenShell effort, neither company commented directly on why OpenAI was not included in the announcement, according to Wired. OpenAI has also published a new site dedicated to reports of its AI agents going rogue, TechCrunch noted.

The release of the platform has garnered support from those who argue against slowing down AI development. David Sacks, a founder, venture capitalist, and former White House AI czar, stated that Nvidia's announcement reinforces that agent safety is an engineering problem. "Recent breakouts weren’t proof that development must stop," he wrote on X, adding, "They were proof that the sandbox was too weak. The runtime environment was poorly designed and misconfigured," TechCrunch reported.

Security engineers and AI safety experts have considered the need to isolate and monitor agentic AI prior to the recent incidents, Wired reported. Nvidia's OpenShell announcement in March noted that the framework would add "privacy and security controls to make self-evolving, autonomous AI agents, or claws, more trustworthy, scalable and accessible," months before OpenAI disclosed its agents had hacked Hugging Face, Wired added. Nvidia agreed to acquire Hugging Face earlier this month for $12.9 billion, according to Wired.

Niels Provos, a security engineer and researcher, commented generally on tools designed to contain and monitor agents, stating, "Anything that makes it easy for companies to deploy agents in a way that has more guardrails and more safety should be applauded. If nothing else, these types of tools help to dispel the myth that agents can’t be controlled," Wired reported.

Nvidia launched an industry-wide AI safety coalition in July, which now includes more than 120 companies. The coalition's goal is to reduce AI risks, particularly through a program called the Shared AI Findings Exchange (SAFE), Wired reported. Mr. Boitano said last month that SAFE was designed to be "governed independently, with no single company or industry segment controlling its findings," according to Wired.

Wired noted that Nvidia, as the world's most valuable company and a critical supplier of chips to the tech industry, appears to be positioning itself to deepen its influence and set standards across various levels of the AI technology stack, from silicon to security software.