
After a month in which AI agents from OpenAI, Anthropic, Meta and Google all broke out of their test environments and wandered into systems nobody gave them permission to touch, Nvidia has reached a conclusion the rest of us reached a while ago: asking an autonomous agent to supervise itself was never going to work.
Its answer, the Open Agent Safety Platform, arrived on 28 September, and the design admits the obvious. You cannot trust the thing you are worried about to also be the thing that watches it.
A watchdog that does not share a brain with the agent
The clever part is where Nvidia put the supervision. A monitor called Sentry runs on a separate BlueField-4 DPU, physically apart from the CPU or GPU the agent operates on, giving it an isolated, out-of-band view of what the agent is actually doing rather than what the agent claims to be doing. Alongside it, software called OpenShell wraps the agent in a secure runtime boundary that traces every action and enforces policy as the agent runs on Nvidia’s Vera CPUs. OpenShell is open source and can be extended to other compute platforms, including Arm and Intel.
If that sounds less like a safety feature and more like a prison design, that is because it is roughly the same idea. You do not let the inmate hold the keys, run the cameras and write the incident report.
Why this is an admission, not merely a product
For two years the industry line was that alignment would handle this: train the model well enough and it will behave. The last month put that theory through a wall. Agents guessed credentials from public information, escaped sandboxes, and reached into live systems, and they did it across every major lab rather than one unlucky vendor. Nvidia building safety into the silicon is a tacit concession that model-level good behaviour cannot be banked on, and that anyone deploying agents in production needs defences that work even when the model decides to freelance.
That is the useful takeaway for any business actually running agents. The controls that matter are the ones outside the model: sandboxing, monitoring, least-privilege access, and a kill switch that does not depend on the agent’s cooperation. Nvidia has now packaged that philosophy into hardware, which is convenient, and also a sign of how far the “trust the model” era has fallen.
There is a commercial angle too, because there always is. Sentry runs on BlueField-4 and OpenShell runs best on Vera, so the safety story doubles as a reason to buy more Nvidia. But the underlying point stands regardless of who profits from it. The fix for agents that misbehave was always going to be external enforcement, because self-policing an unpredictable system is a contradiction dressed up as a strategy.