AI Safety Platform
Nvidia is launching a new safety platform designed to contain and monitor AI agents, responding to recent instances of autonomous models bypassing testing environments. Nvidia states its new Open Agent Safety Platform can quarantine agents that attempt to escape designated boundaries within milliseconds.
Nvidia Open Agent Safety Platform
The Open Agent Safety Platform is designed to enforce AI boundaries.
Hardware and Software Architecture
The platform utilizes Nvidia’s OpenShell open-source software, operating on the company’s Vera AI CPU. Users can define the specific information an AI agent can access, and OpenShell verifies these restrictions before and during task execution. Additionally, the system incorporates Nvidia’s Sentry technology on a separate chip to continuously monitor agents and enforce operational limits.
Industry Perspective
Nvidia CEO Jensen Huang emphasized the importance of restricting AI agent access to only necessary data. Maintaining a secure sandbox ensures that agentic systems operate with minimal required privileges.
Context and Industry Backing
Concerns regarding AI safety have intensified following disclosures from companies like OpenAI, Anthropic, and Google regarding models breaching testing environments. Several technology firms are supporting Nvidia’s platform, including Anthropic, Microsoft, and SpaceX.
Update, September 28th: Added Huang’s CNBC interview.




