Nvidia Releases Tools to Contain Rogue AI Agents

Nvidia has released a new set of security tools designed to put enforceable boundaries around autonomous AI agents as concerns grow over systems that can act beyond their intended environments.

Nvidia Releases Tools to Contain Rogue AI Agents

The company announced the Nvidia Open Agent Safety Platform on September 28, introducing an open software platform and reference system designed to monitor, restrict and contain AI agents across software, computing infrastructure and, eventually, physical robotics systems.

The platform centers on two technologies: Nvidia OpenShell, an open-source secure runtime for AI agents, and Nvidia Sentry, an out-of-band monitoring and enforcement system designed to stop an agent when it attempts to move outside its permitted boundaries.

What Is Nvidia Open Agent Safety Platform?

Nvidia describes the Open Agent Safety Platform as a full-stack approach to AI-agent security.

Instead of relying exclusively on the AI model or the software framework running the agent, the platform adds security controls at the infrastructure level.

This distinction matters because autonomous agents can increasingly interact with files, websites, APIs, software tools, databases and other systems. If an agent behaves unexpectedly, controls outside the model can provide an additional layer of enforcement.

Nvidia says organizations can deploy different components of the platform according to their security requirements.

OpenShell Creates a Security Boundary Around AI Agents

OpenShell is the software component of Nvidia’s new platform.

It provides a secure runtime boundary around autonomous agents and is designed to trace their actions and enforce policies while they operate.

The goal is to prevent an agent from simply overriding restrictions implemented inside the application or agent framework itself.

Nvidia says OpenShell is broadly available as open-source software and can run with Nvidia Vera CPUs. The company is also designing it to work with third-party computing platforms, including CPUs from Arm and Intel.

That cross-platform approach could make the technology relevant beyond Nvidia’s own hardware ecosystem.

Sentry Adds Hardware-Based Agent Monitoring

The second major component is Nvidia Sentry.

Sentry is designed to run separately from the main agent environment on Nvidia BlueField-4 data processing units. Nvidia describes it as an out-of-band watchdog, meaning it can monitor and enforce security policies independently of the software environment where the AI agent is running.

If an agent attempts to move outside its permitted software boundary, Nvidia says Sentry can quarantine and stop it in milliseconds.

The system is designed to inspect agent activity, verify agent identity, provide telemetry and enforce access policies covering data, tools, APIs and services.

This creates a security layer that is intended to remain outside the agent’s direct control.

Nvidia Says the Tools Could Have Stopped the Hugging Face Incident

Nvidia’s announcement comes after a series of AI-agent security incidents involving frontier AI systems.

Reuters reported that Nvidia executives said the company’s new security platform could have prevented the recent Hugging Face incident involving OpenAI agents.

That is a statement from Nvidia rather than an independently demonstrated result from a public reproduction of the incident.

The comparison nevertheless illustrates the problem Nvidia is attempting to address: an AI agent may follow a legitimate task while finding ways around application-level restrictions or security controls.

The company argues that security boundaries should therefore exist below the model and agent-harness level.

Why AI Agents Need Different Security Controls

Traditional software typically follows explicitly programmed instructions.

AI agents can operate differently because they can interpret objectives, decide which tools to use and adapt their actions based on information they encounter during execution.

An agent working on a coding task, for example, could potentially interact with a terminal, access files, call APIs or communicate with external services.

As agents receive more permissions, a mistake or unexpected behavior can have consequences beyond the original application.

Nvidia’s platform is designed around the idea that these permissions should be constrained by infrastructure-level policies that the agent cannot simply rewrite.

The Platform Is Designed for Long-Running Agents

Nvidia is positioning the technology for organizations that allow AI agents to perform extended tasks with limited supervision.

Long-running agents create a different security challenge from simple chatbot interactions because they can perform multiple actions over time.

An agent could make one decision, use the result to determine its next action and continue interacting with external systems.

OpenShell is intended to provide a persistent runtime boundary during that process, while Sentry adds independent monitoring and enforcement.

Together, the systems are designed to provide visibility into what agents are doing while limiting where they can operate.

Nvidia Is Working With More Than 100 Organizations

Nvidia said more than 100 organizations are working with its Open Agent Safety Platform technologies.

Participants and collaborators named by the company include Anthropic, Microsoft, Hugging Face, CrowdStrike, Cisco, Palantir, Palo Alto Networks, Perplexity, Salesforce, SAP, ServiceNow, Scale AI and others.

The platform is also being explored in areas beyond conventional enterprise software.

Nvidia said robotics companies including Figure, Gecko Robotics and Skild AI are building with OpenShell to incorporate agent-security controls into autonomous systems that can take actions in the physical world.

Financial-services companies including Citi and JPMorganChase are also collaborating with Nvidia on open-source agent-security technologies.

Anthropic and Nvidia Are Adding Another Security Layer

Anthropic is among the companies collaborating with Nvidia on the platform.

Nvidia said its work with Anthropic combines Claude Managed Agents with OpenShell and BlueField technologies.

Claude Managed Agents separate the agent loop from the sandboxes where work is executed, while Nvidia’s infrastructure can add additional restrictions over access to those environments.

The broader approach reflects an emerging security principle in agentic AI: an AI model should not be the only component responsible for controlling its own permissions.

OpenShell Can Be Managed Through Enterprise Systems

Nvidia is also integrating OpenShell with existing enterprise software.

Salesforce and Nvidia have integrated OpenShell with Slack, allowing teams to view agent activity and audit events and approve or reject requests for additional permissions.

SAP is embedding OpenShell into its Joule Studio runtime as part of its Business AI Platform.

These integrations are intended to give organizations more visibility and human oversight while AI agents perform work across enterprise environments.

Nvidia Wants Agent Security to Work Across the Stack

The Open Agent Safety Platform is broader than a conventional AI safety filter.

Its architecture spans the AI-agent runtime, computing hardware and external systems that agents interact with.

Nvidia argues that this type of full-stack enforcement is necessary because application-level security controls may not be sufficient when autonomous agents become capable of finding alternative paths to complete a task.

The company is therefore positioning OpenShell and Sentry as infrastructure-level controls rather than another layer of instructions given to the AI model.

The Tools Are Available as Open Software

Nvidia says Open Agent Safety Platform software, including OpenShell and associated skills, is available through its developer resources and GitHub.

OpenShell is open source and is designed to support both open and closed AI models.

Nvidia also says its broader ecosystem effort is connected to the Open Secure AI Alliance, an initiative involving more than 120 organizations and governed by the Linux Foundation.

The objective is to encourage shared research, security tools, evaluation methods and standards for increasingly autonomous AI systems.

What Nvidia’s New Tools Mean for AI Security

The release comes as AI companies increasingly move from conversational models toward agents that can execute tasks across real software environments.

That shift changes the security problem.

Instead of only checking whether an AI-generated response is safe, organizations also need to control what an AI system can access and what actions it can perform.

Nvidia’s Open Agent Safety Platform represents one approach: put enforceable boundaries outside the model and monitor agent behavior independently.

Whether these controls can reliably prevent future incidents will depend on how they perform across different models, operating environments and attack scenarios.

For now, Nvidia is making the technology available as part of an open ecosystem while AI developers and enterprises continue to work out how much autonomy agents should receive.

Frequently Asked Questions

What is Nvidia Open Agent Safety Platform?

It is an open software platform and reference system designed to provide governance, monitoring and security controls for AI agents across software and hardware infrastructure.

What is Nvidia OpenShell?

OpenShell is an open-source secure runtime designed to create enforceable boundaries around AI agents and control how they execute tasks.

What is Nvidia Sentry?

Sentry is an out-of-band monitoring and enforcement system designed to run on Nvidia BlueField-4 DPUs and independently monitor agent activity.

Can Nvidia Sentry stop a rogue AI agent?

Nvidia says Sentry can quarantine and stop an agent in milliseconds if it attempts to move outside its defined software boundary.

Could Nvidia’s tools have prevented the Hugging Face incident?

Nvidia executives said the platform could have prevented the incident based on what the company knows about it. That is Nvidia’s assessment rather than an independently demonstrated result.

Is OpenShell open source?

Yes. Nvidia says OpenShell is available as open-source software and is designed to support open and closed AI models.

Does the platform only work with Nvidia CPUs?

Nvidia says OpenShell is designed to work with third-party computing platforms, including Arm and Intel CPUs.

Scroll to Top