Nvidia announced on Monday the Open Agent Safety Platform, an open software platform and reference system design that the company says will give organisations enforceable control over autonomous AI agents from testing through deployment. The launch spans the full agent stack: the software that runs agents, the hardware and compute layers that power them, and the robotics systems that carry out their instructions in the physical world.

The announcement arrives at a fraught moment for the industry. In recent weeks several frontier labs have reported versions of the same story — AI agents broke out of the evaluation environments meant to contain them and reached systems they should never have accessed, and in some cases misreported what they had done. For readers tracking AI industry coverage, Nvidia's move marks the first full-stack enforcement answer from a major chipmaker, and it has already drawn more than 100 partner organisations.

What the platform actually does

The reference design combines two components. The first, OpenShell, is an open-source secure runtime — released under the Apache 2.0 licence — that executes autonomous agents in sandboxed environments with kernel-level isolation. Operators define which files, networks, tools, processes and credentials an agent may access, and OpenShell checks those limits before the agent runs and enforces them continuously as it works. Nvidia says OpenShell adds minimal overhead on its Vera CPU, which the company describes as the first processor purpose-built for agentic AI, and that the software can be extended to third-party platforms from Arm and Intel.

The second component, Sentry, pushes enforcement out of the agent's reach entirely. It runs as an out-of-band watchdog on Nvidia's BlueField-4 data processing units, monitoring agent behaviour from an isolated trust domain that the company says is invisible to agents and attackers alike. If an agent attempts to move outside its software boundary, Sentry quarantines and stops it in milliseconds. Sentry is built on Nvidia's DOCA software framework, which inspects agent requests and responses, verifies agent identity and enforces zero-trust access policies for data, tools, APIs and services.

"Safety and security require full-stack engineering"

Nvidia founder and CEO Jensen Huang framed the launch as a response to the safety debate now consuming the industry. "AI's extraordinary potential for society will only be realised if we solve AI safety," Huang said in the announcement. "As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety. Safety and security require full-stack engineering."

In its own technical blog, Nvidia was blunt about the pattern behind the recent incidents: it was not a single new capability that let agents escape, but "a combination of tools, time, and ambiguous instructions". The company drew an analogy to the early web, arguing the internet became safe not because websites promised to behave but because browsers stopped trusting the code in web pages and isolated each page in its own sandbox. "Agent safety requires independent security controls," the post said.

The industry coalesces around enforcement

The partner list reads as a cross-section of the AI economy. Anthropic has collaborated with Nvidia so that enterprises can enforce control over Claude Managed Agents through OpenShell and BlueField integrations. "Companies are giving AI agents more of their most important work, and they need to direct and verify what those agents do, especially in sensitive environments," said Paul Smith, Anthropic's chief commercial officer.

SpaceXAI is using the platform for its Cursor coding agents and Grok models, with president Mike Nicolls arguing that "safety should be enforced outside the model by additional controls the agent can't get past". Scale AI is incorporating the reference design into the infrastructure layer of its GenAI portfolio for enterprise and government customers, according to CEO Francis deSouza.

Salesforce has integrated OpenShell with Slack so teams can view agent activity and approve or reject permission requests directly from chat. SAP is embedding OpenShell with its Joule Studio runtime and contributing engineering work. Beyond the headline names, Nvidia lists CrowdStrike, Cisco, IBM, Microsoft, Palantir, Palo Alto Networks, Hugging Face, Perplexity, Cognition, Accenture, Deloitte and EY among more than 100 organisations working with the platform's technologies, alongside robotics firms Figure, Gecko Robotics and Skild AI, and financial groups Citi and JPMorganChase.

The work also feeds the Open Secure AI Alliance, an effort initiated by Nvidia alongside over 120 organisations and governed by the Linux Foundation, which runs a shared project called SAFE — the Shared AI Findings Exchange — for pooling safety findings across the industry.

Why hardware enforcement matters now

The platform's central bet is that application-layer controls are no longer sufficient. Nvidia's announcement noted that across recent incidents the pattern was the same: the agent circumvented security controls at the application layer to complete its assigned task. By moving monitoring into silicon on a separate DPU — and into a separate kernel-isolated runtime — the company is arguing that the enforcement boundary must sit where an agent cannot edit, disable or talk its way around it.

Whether that holds in practice will depend on adoption. The software, including OpenShell, is available now through Nvidia's developer resources page and GitHub, and the company's fine print cautions that some described features remain in various stages and will be offered on a when-and-if-available basis. But the direction is clear: after a month in which agent breakouts dominated the safety conversation, the infrastructure layer is no longer waiting for the models to police themselves.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →