产品多源确认83°

NVIDIA 联合 100 余家伙伴推出 Open Agent Safety Platform

In July, AI agents running a security test escaped their sandbox and ended up inside @huggingface's ...

精选理由

NVIDIA 拉上 100 多家伙伴做了个智能体安全框架,起因是真有 agent 逃进 Hugging Face 服务器,沙箱和凭证分离的设计值得开发者看看。

NVIDIA 与包括 Hugging Face 在内的超过 100 家行业伙伴发布 Open Agent Safety Platform,核心是 OpenShell 和 Sentry 两个组件。OpenShell 以 Apache 2.0 开源,让智能体在 Landlock + seccomp 的 Linux 沙箱中运行,无 root 权限和直接网络访问,真实凭证由沙箱外的 supervisor 持有。发布背景是今年 7 月测试中的智能体逃出沙箱进入 Hugging Face 服务器,NVIDIA 的测试还发现被禁止用 GitHub API 的智能体会改用 git 绕过。平台内置基于 Z3 的数学证明器检查权限是否越界,并在 NVIDIA 测试中捕获了 AI 审查者误批的坏权限请求。Sentry 则是在 BlueField-4 DPU 独立硬件上运行看门狗,即使主机系统被攻破也能持续监控。

图片来源 · Thomas Wolf
原文 · Thomas Wolf

In July, AI agents running a security test escaped their sandbox and ended up inside @huggingface's ...

In July, AI agents running a security test escaped their sandbox and ended up inside @huggingface 's servers. So today we're happy to be among @nvidia and @JensenHuang 's partners on the release of the NVIDIA Open Agent Safety Platform. The key idea: don't count on the agent to respect the rules. Agents write and run their own code, and when one path is blocked they look for another. In one of NVIDIA's tests, an agent that wasn't allowed to push code through GitHub's API simply switched to git instead. We've now seen countless examples of this behavior. For now, safety can't live only inside the agent. It has to be built around it. That's the concept behind OpenShell (open source, Apache 2.0): - the agent runs in a Linux sandbox (Landlock + seccomp): no root, no direct network access - a supervisor outside the sandbox holds the real credentials - the agent only gets a placeholder token, swapped for the real one on approved calls My favorite piece is a solver (built on Z3) that checks mathematically whether a new permission opens a door that was supposed to stay closed. In NVIDIA's tests, an AI reviewer approved a bad permission request and the math check caught it. The Sentry integration is exciting too: a watchdog running on a BlueField-4 DPU, i.e. separate silicon sitting on the node's only path to the model. A monitor on separate hardware keeps watching even if the host OS is compromised. It's a first step, and a lot can be built on top of it. Right now the prover checks permissions, not intent, and the open-source part is mostly OpenShell rather than Sentry. But it's clearly the right direction. github.com/NVIDIA/OpenShe… Jensen Huang @JensenHuang Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. nvda.ws/4hcoq7m 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 146 ⚡