Nvidia has released a set of software tools designed to improve the safety of AI agents, saying the technology could have prevented a recent hack involving Hugging Face, the AI coding platform it acquired for nearly $13 billion.
The move comes as OpenAI and Anthropic investigate several incidents involving AI agents that gained access to commercial and government systems while carrying out complex tasks.
One of Nvidia’s new tools, OpenShell, uses hardware features in Nvidia processors to isolate and control AI agents. The company said it is working with Arm Holdings and Intel to make the system compatible with their processors.
Nvidia is launching the safety tools with dozens of partners, including Anthropic.
Justin Boitano, Nvidia’s vice president and general manager of enterprise computing, said the technology could have prevented the Hugging Face breach if it had been used during early model evaluations.
Another tool, Sentry, works with OpenShell and uses a separate Nvidia chip to stop an AI agent if it attempts to escape its protected environment.
Nvidia said its systems use mathematical techniques to identify attempts by AI agents to bypass safeguards, including efforts to create additional agents to circumvent restrictions.
Leave a comment