Nvidia Launches Open Agent Safety Platform for AI Agents
Nvidia Open Agent Safety Platform is bringing a two-layer security model to autonomous AI agents, combining an open-source runtime with a separate hardware watchdog. Announced on September 28, the system is designed to constrain agents before deployment and interrupt them when their behavior crosses an approved boundary.
The release turns Nvidia's March preview of OpenShell into a broadly available product and adds Sentry, which runs independently on a BlueField-4 data-processing unit. Nvidia says the reference design can isolate agents, monitor their activity continuously and quarantine violations in milliseconds.
The platform has three central components:
- OpenShell sets file, network and credential permissions.
- Sentry watches agent behavior outside the host system.
- A reference design connects Vera CPUs with BlueField-4 DPUs.
Nvidia Open Agent Safety Platform Combines Two Enforcement Layers
The platform separates prevention from response. OpenShell creates a restricted execution environment around an agent, while Sentry observes the resulting activity from a different processor. That division is intended to keep a compromised agent from weakening the same controls that are supposed to contain it.
According to Nvidia's announcement, OpenShell can be used during testing and production without being tied to one model or agent framework. Sentry adds continuous enforcement for organizations that want the monitoring layer physically separated from the compute running the agent.
The accompanying reference architecture places OpenShell on Nvidia Vera CPUs and Sentry on BlueField-4 DPUs. Nvidia presents the design as a starting point for data centers, cloud services and enterprise systems rather than as a mandatory hardware configuration for every OpenShell deployment.
How OpenShell Restricts Files, Credentials and Networks
OpenShell provides sandboxed environments with kernel-level isolation. Administrators express permissions in declarative YAML policies, defining which files an agent can read or change, which network destinations it can reach and which credentials it may use. The default approach is least privilege: access must be granted rather than assumed.
Those controls matter because an agent can perform long chains of actions without pausing for human approval. A coding agent, for example, may inspect repositories, install packages, call external services and execute commands. A model-level refusal is not a reliable security boundary if a prompt injection or software flaw changes the agent's behavior.
OpenShell is available under the Apache 2.0 license, and Nvidia says it can be extended beyond its own processors. Reuters reported that Nvidia is working with Arm and Intel on compatible implementations, widening the potential market beyond systems built entirely around Nvidia hardware.
Open source gives security teams access to the enforcement code and policy model, but it does not make deployment automatic. Organizations still need accurate inventories, narrow permissions, protected secrets and tests that verify legitimate agent work continues after restrictions are applied.
Sentry Adds an Independent BlueField-4 Watchdog
Sentry is the new element in the September release. It runs on the BlueField-4 DPU and watches the host from outside the agent's execution environment. Nvidia says it can detect an attempt to cross established boundaries and quarantine or stop the offending process in milliseconds.
That out-of-band design addresses a familiar security problem: software should not be solely responsible for judging its own conduct. If an attacker compromises the host or an agent spawns subordinate processes, a monitoring component on the same system may lose visibility or authority. Sentry is intended to retain both.
Nvidia has also argued that the platform could have prevented an earlier compromise involving a Hugging Face coding agent. That is a retrospective vendor claim, not an independently reproduced result. The incident illustrates the threat model, but buyers will still need evidence about detection accuracy, false positives, workload overhead and recovery procedures.
Hardware separation also raises practical questions. BlueField-4 availability, integration work and operating cost will shape where Sentry is used first. High-risk coding, infrastructure and research agents are more plausible early targets than low-permission assistants that cannot reach sensitive systems.
Partners Test Nvidia's Agent-Security Approach
Nvidia says more than 100 companies are working with the platform, including Anthropic, Cisco, CrowdStrike, Dell, HPE, Microsoft, Red Hat, Salesforce, SAP and ServiceNow. The list spans model developers, security vendors, server manufacturers and enterprise software providers, reflecting the number of layers involved in agent deployment.
SpaceXAI is using the platform with Cursor coding agents and Grok models, according to Nvidia. Anthropic said hardware-isolated enforcement can complement model-level safeguards by giving enterprises more control over how agents interact with files, networks and credentials.
Participation does not mean every partner has completed a production rollout. The important near-term measure will be whether OpenShell policies and Sentry signals integrate with existing identity, logging and incident-response systems. Security teams need controls they can audit and operate, not another isolated alert stream.
Nvidia's release arrives as autonomous agents gain broader access to code, business applications and cloud infrastructure. The platform's central bet is that agent safety requires enforceable system boundaries alongside model alignment. Whether it becomes a common standard will depend on interoperability and transparent performance data, not the size of its launch partner list.
Further Reading