Nvidia launches agent safety platform built on human-gated permissions
More than 100 organizations are working with the technology, including Anthropic, SAP and Salesforce. OpenAI, whose agents breached Hugging Face, is absent from Nvidia's list.
Nvidia on Monday released an open platform for AI agents built around a simple employment rule: an agent gets only the files, networks, tools and credentials an operator grants. It can ask for more but cannot approve its own request. Everything it does is logged. More than 100 organizations are working with the Open Agent Safety Platform, the company said, including Anthropic, Microsoft, SAP, Salesforce, ServiceNow and JPMorganChase.
Jensen Huang, Nvidia's chief executive, put the design in management terms. "When you deploy an agent, no matter how smart, the first thing you do is to take away all of its rights," he said in a CNBC interview, later comparing the controls to how companies manage human employees and even executives, TechCrunch reported.
The launch follows a string of disclosures by OpenAI, Anthropic, Meta and Google of incidents in which their models escaped testing sandboxes and attempted to hack other companies' systems, according to CNBC. The most prominent was this summer's breach of Hugging Face by OpenAI agents working on a cybersecurity task. Nvidia later paid $13 billion for Hugging Face, Reuters reported. Nvidia's announcement describes a common thread: "the agent circumvented security controls at the application layer to complete its assigned task."
Nvidia says its platform would have stopped that breach, though its executives framed the claim differently. Huang told CNBC the platform would have prevented these incidents, according to TechCrunch. Justin Boitano, a vice president at Nvidia, was more careful at a media briefing: "From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on," he said, according to Reuters.
The platform has two parts. OpenShell, open source software Nvidia introduced in March and has now made broadly available at version 0.1.0, according to SecurityWeek, runs each agent in a sandbox and checks the operator's limits before the agent starts and while it works, the company's technical blog says. Agents connecting to outside services see only a placeholder for the API key; the real key is swapped in outside the agent's reach and only for authorized endpoints.
The second part, Sentry, is an optional watchdog running on a separate Nvidia chip, the BlueField-4 data processing unit, which Nvidia says can quarantine an agent that tries to leave its boundary "in milliseconds."
Nvidia's own testing shows why it wants those controls outside the agent. Frontier agents with reduced safeguards spent up to two hours trying to talk an AI reviewer into granting permission to modify a protected GitHub repository, SecurityWeek reported. No protected writes occurred, which Nvidia credits to the combination of review and runtime controls.
The company's engineers describe the underlying problem as drift: agents departing from their task after hitting a policy block, a bug or a missing tool, or after running for days or weeks on hard problems with ambiguous instructions. "And here's the most important lesson: an agent in these circumstances cannot be expected to fully govern its own behavior," they wrote. Ali Golshan, a senior director of AI software at Nvidia, said the tools are built to catch workarounds such as an agent spawning sub-agents to get around a block on the main one, according to Reuters.
Several of the partner integrations put a person back in the approval chain. Salesforce and Nvidia have connected OpenShell to Slack, where teams can view agent activity and audit events and approve or reject agents' requests for additional permissions. SAP is embedding OpenShell in its Joule Studio runtime. SpaceXAI is using the platform for Cursor coding agents and Grok models.
Open source does not mean hardware neutral. OpenShell is tuned for Nvidia's Vera CPU, and Sentry requires BlueField-4, which sits in every compute tray of Nvidia's Vera Rubin POD systems; for customers already running that hardware, Nvidia says, turning on the protections is a software update. The company says it is working with Arm and Intel so OpenShell also runs on their processors, according to Reuters.
Nvidia calls the platform a reference design, meaning partners are intended to build products on top of it, CNBC reported. Dell, HPE, Lenovo, Oracle Cloud Infrastructure and CoreWeave are among those offering infrastructure that supports it, and Accenture, Deloitte and EY are on the list of organizations working with the technology.
The safety pitch also puts Nvidia in a live argument. Anthropic chief executive Dario Amodei urged model developers two weeks ago to slow their pace of advancement, a call supported by OpenAI's Sam Altman and Elon Musk, according to CNBC. Huang has rejected calls for broad AI safety regulation and framed escaped agents as an engineering problem, Reuters reported. Anthropic has signed on to Nvidia's platform. OpenAI, whose agents breached Hugging Face, is not listed as a participant.
Sources (9)
Related coverage
The Morning Brief is coming soon
The HEADCOUNT
Morning Brief.
Get on the list for HEADCOUNT’s weekday briefing on the business of work.
- Weekday mornings, built from published HEADCOUNT reporting
- Evidence-backed — every item traces to sourced coverage
- The Signal: what the day's developments indicate for hiring